# Evaluate your server with models

Run an evaluation of your registered MCP server in the dashboard, and read it: how the cases are written, the distractor servers, what the metrics mean, and how the units are counted.

The [linter](https://forecall.dev/en/docs/reading-scores) scores how your tools are written. An evaluation checks what
models do with them: for utterances written for each of your tools, it asks models which tool they
would call, and checks the arguments they give. Evaluations come with the Pro and Team plans. Run
them in the dashboard as below, or [with the REST API](https://forecall.dev/en/docs/api#evaluations).

## Start an evaluation

Open **Evaluations** in the dashboard at [app.forecall.dev](https://app.forecall.dev) and **Start
an evaluation**. Choose one of your [registered servers](https://forecall.dev/en/docs/dashboard): its latest version is
evaluated. Choose the models, any distractor servers and the utterances' language, then
**Estimate**. Nothing is taken until you start; the button says how many units it takes.

The models you can choose are `anthropic/claude-opus-5-5`, `anthropic/claude-sonnet-5-5`,
`openai/gpt-6-astra` and `openai/gpt-6.1-sol`; `google/gemini-3.8-flash` is paused for now. The
middle ones, `anthropic/claude-sonnet-5-5` and `openai/gpt-6.1-sol`, are chosen at first.

An evaluation takes a few minutes. Its page updates by itself while it runs (reload it when
JavaScript is off), and shows each model's results as soon as its run is done.

## The cases

A model (`anthropic/claude-sonnet-5-5`) writes the cases from your tools' names, descriptions and
input schemas. There are three kinds:

- **For the tool**: utterances that call for one tool, 8 for each tool on Pro and 20 on Team.
- **Easily confused**: for each pair of tools the linter finds
  [confusable](https://forecall.dev/en/docs/reading-scores), two utterances for each of the two tools.
- **Calls for no tool**: utterances none of your tools should answer, a fifth as many as those for
  the tools.

The utterances are in English, or in Japanese on Team. The cases are kept with the evaluation: a
later evaluation of the same version with the same settings uses the same cases, so the two compare
fairly, and writing them is not counted again.

## Distractor servers

Your tools are seldom the only ones a model sees. Distractor servers put the tools of public
servers beside yours: `bhived`, `context7`, `debugbase`, `filesystem`, `firecrawl`, `memory`,
`notion` and `playwright`. None are chosen at first. A call to a distractor's tool counts as another
tool.

The estimate warns you when a distractor has a tool named like one of yours, as a call to either
cannot be told apart, and when the models see more than 60 tools in all, which they handle less
well.

## Read the results

For each model, the evaluation's page shows three metrics, each with the counts behind it:

| Metric | What it is |
|---|---|
| Tool choice (`selectionAcc`) | Of the cases written for a tool, the share that called that tool. |
| Arguments (`argValidity`) | Of those calls, the share whose arguments fit the tool's input schema. A schema that cannot be read is left out. |
| False calls (`falseCallRate`) | Of the cases that call for no tool, the share that got a call. Lower is better. |

An answer cut off at the length limit is left out of all three. A metric with nothing to count
shows a dash.

The page compares each model with the server's previous evaluation that has results, and shows
**Which tool was called**: for each model, a row for the tool each case was written for and a
column for the tool it called, the diagonal being right. Past 12 tools, it shows only the rows
with a mistake and the tools called by mistake, and **Show all** brings back the rest. **See
each case** lists every utterance with the tool each model called and a verdict, narrowed if you
like to the cases any model missed, those one model missed, or one kind of mistake. A call with
wrong arguments says which argument and how: missing, of the wrong type, not an allowed value.
The utterances come from your tools, so only your organization's members see them.

## Units

Each model's run takes units by the model's price and the size of what it is asked: the tools it is
shown, yours and the distractors', and each utterance. Writing the cases takes units too, counted
with the first model. A result your organization already has, for the same model, tools and
utterance, is reused and not counted again. The estimate shows the total, each part, and this
month's units left with your credits.

An evaluation takes its units from the month's allowance first, then from your credits, the ones
that expire first. When both fall short, it does not start, and neither does it while the
organization's last payment has failed, until it is paid. If it fails, the units of the cases it
did not finish come back as credits, valid for six months.
