Evaluate your server with models
Run an evaluation of your registered MCP server in the dashboard, and read it: how the cases are written, the distractor servers, what the metrics mean, and how the units are counted.
On this page
The linter scores how your tools are written. An evaluation checks what models do with them: for utterances written for each of your tools, it asks models which tool they would call, and checks the arguments they give. Evaluations come with the Pro and Team plans. Run them in the dashboard as below, or with the REST API.
Start an evaluation
Open Evaluations in the dashboard at app.forecall.dev and Start an evaluation. Choose one of your registered servers: its latest version is evaluated. Choose the models, any distractor servers and the utterances' language, then Estimate. Nothing is taken until you start; the button says how many units it takes.
The models you can choose are anthropic/claude-opus-5-5, anthropic/claude-sonnet-5-5,
openai/gpt-6-astra and openai/gpt-6.1-sol; google/gemini-3.8-flash is paused for now. The
middle ones, anthropic/claude-sonnet-5-5 and openai/gpt-6.1-sol, are chosen at first.
An evaluation takes a few minutes. Its page updates by itself while it runs (reload it when JavaScript is off), and shows each model's results as soon as its run is done.
The cases
A model (anthropic/claude-sonnet-5-5) writes the cases from your tools' names, descriptions and
input schemas. There are three kinds:
- For the tool: utterances that call for one tool, 8 for each tool on Pro and 20 on Team.
- Easily confused: for each pair of tools the linter finds confusable, two utterances for each of the two tools.
- Calls for no tool: utterances none of your tools should answer, a fifth as many as those for the tools.
The utterances are in English, or in Japanese on Team. The cases are kept with the evaluation: a later evaluation of the same version with the same settings uses the same cases, so the two compare fairly, and writing them is not counted again.
Distractor servers
Your tools are seldom the only ones a model sees. Distractor servers put the tools of public
servers beside yours: bhived, context7, debugbase, filesystem, firecrawl, memory,
notion and playwright. None are chosen at first. A call to a distractor's tool counts as another
tool.
The estimate warns you when a distractor has a tool named like one of yours, as a call to either cannot be told apart, and when the models see more than 60 tools in all, which they handle less well.
Read the results
For each model, the evaluation's page shows three metrics, each with the counts behind it:
| Metric | What it is |
|---|---|
Tool choice (selectionAcc) |
Of the cases written for a tool, the share that called that tool. |
Arguments (argValidity) |
Of those calls, the share whose arguments fit the tool's input schema. A schema that cannot be read is left out. |
False calls (falseCallRate) |
Of the cases that call for no tool, the share that got a call. Lower is better. |
An answer cut off at the length limit is left out of all three. A metric with nothing to count shows a dash.
The page compares each model with the server's previous evaluation that has results, and shows Which tool was called: for each model, a row for the tool each case was written for and a column for the tool it called, the diagonal being right. Past 12 tools, it shows only the rows with a mistake and the tools called by mistake, and Show all brings back the rest. See each case lists every utterance with the tool each model called and a verdict, narrowed if you like to the cases any model missed, those one model missed, or one kind of mistake. A call with wrong arguments says which argument and how: missing, of the wrong type, not an allowed value. The utterances come from your tools, so only your organization's members see them.
Units
Each model's run takes units by the model's price and the size of what it is asked: the tools it is shown, yours and the distractors', and each utterance. Writing the cases takes units too, counted with the first model. A result your organization already has, for the same model, tools and utterance, is reused and not counted again. The estimate shows the total, each part, and this month's units left with your credits.
An evaluation takes its units from the month's allowance first, then from your credits, the ones that expire first. When both fall short, it does not start, and neither does it while the organization's last payment has failed, until it is paid. If it fails, the units of the cases it did not finish come back as credits, valid for six months.