# Know before you call.

Static linter · free · no sign-up

Forecall measures whether AI agents choose and call your MCP tools correctly, and shows you what to fix.

[Score your tool descriptions for free](https://forecall.dev/en/lint): Paste your tools/list JSON. No sign-up needed.

[Forecall on Product Hunt](https://www.producthunt.com/products/forecall?embed=true&utm_source=badge-featured&utm_medium=badge&utm_campaign=badge-forecall) · [Forecall on Orynth](https://www.orynth.dev/projects/forecall)

Sample result · Notion's official server:

```text
$ forecall lint notion-tools.json
24 tools · scoring rules v1
 20  API-query-data-source
 24  API-create-a-comment
 24  API-update-a-data-source
… and 21 more tools
average score 33.0 · confusable pairs 57
```

## What we found in public MCP servers

We scored 8 well-known MCP servers, 126 tools in all. Their average scores ranged from 33 to 70 out of 100. Notion's official server alone has 57 pairs of tools that an AI can easily mix up.

- Servers scored: 8
- Tools: 126
- Range of average scores: 33–70
- Confusable pairs in Notion's server alone: 57

| Server | Tools | Average score | Confusable pairs |
|---|---:|---:|---:|
| Notion API | 24 | 33.0 (Rain) | 57 |
| Playwright | 25 | 41.9 (Cloudy) | 0 |
| bhived-mcp | 12 | 70.3 (Clear) | 0 |

Three of the eight servers, scored with the static linter (scoring rules v1). 70 and above: clear · 40–69: cloudy · below 40: rain

## Beyond the linter

- **Evaluations**: Real models choose and call your tools, and Forecall measures how often they get it right. On the Pro and Team plans. [How evaluations work](https://forecall.dev/en/docs/evaluations)
- **Failure KB**: Known failures of MCP tool calls and their workarounds, for your agent to look up before and after a call. A record is verified only after agents of two other teams reproduce it. [Connect your agent](https://forecall.dev/en/docs/kb)
