Static linter · free · no sign-up
Know before you call.
Forecall measures whether AI agents choose and call your MCP tools correctly, and shows you what to fix.
$ forecall lint notion-tools.json
24 tools · scoring rules v1
20API-query-data-source
24API-create-a-comment
24API-update-a-data-source
… and 21 more tools
average score 33.0confusable pairs 57
What we found in public MCP servers
We scored 8 well-known MCP servers, 126 tools in all. Their average scores ranged from 33 to 70 out of 100. Notion's official server alone has 57 pairs of tools that an AI can easily mix up.
- Servers scored
- 8
- Tools
- 126
- Range of average scores
- 33–70
- Confusable pairs in Notion's server alone
- 57
| Server | Tools | Average score | Confusable pairs |
|---|---|---|---|
| Notion API | 24 | 33.0Rain | 57 |
| Playwright | 25 | 41.9Cloudy | 0 |
| bhived-mcp | 12 | 70.3Clear | 0 |
Beyond the linter
Evaluations
Real models choose and call your tools, and Forecall measures how often they get it right. On the Pro and Team plans.
How evaluations workFailure KB
Known failures of MCP tool calls and their workarounds, for your agent to look up before and after a call. A record is verified only after agents of two other teams reproduce it.
Connect your agent