Getting started
What Forecall measures, what you can do for free, including the Failure KB for your agents, what the paid plans add, and what static scoring cannot tell.
On this page
An AI agent chooses a tool by reading its name, its description and its input schema. When the text is vague, the agent picks the wrong tool, fills in the wrong arguments, or never calls the tool at all. Forecall measures how well your MCP tools are written for that reader, and shows you what to fix.
What Forecall measures
Forecall reads what your MCP server returns for tools/list and scores each tool out of 100 on
six things a model needs to choose and call it:
- its purpose
- when to use it, and when not to
- its arguments
- its return value
- its constraints and side effects
- examples
It also looks across the whole server: too many tools, tools that are easy to mix up, and tools that share the same description. Each finding comes with an issue code and a severity (critical, major or minor), so you can fix the worst first. Read your scores goes through the items and every issue code.
This is static scoring: Forecall reads the text and never runs your server or a model. Measuring how often real models choose and call your tools correctly is the job of evaluations.
Free and paid
Free, with no sign-up:
- The linter on the web: paste your
tools/listJSON at the linter. You get a result page whose URL you can share. Anyone with the URL can see it, search engines do not list it, and it is deleted after 30 days. - The CLI:
npx forecall lint tools.jsonscores a file on your own machine with the same rules and sends nothing anywhere. Use it in CI to keep scores from dropping.
Free, with an account at app.forecall.dev:
- Servers and history: register your servers and follow their scores over time.
- API keys: score from your own scripts and CI through the REST API, within a small monthly allowance.
- The Failure KB: connect your AI agent with
npx forecall setup, so that it looks up known failures of MCP tools and their workarounds before and after a call, within a monthly allowance of lookups.
Paid, on the Pro and Team plans:
- Evaluations: real models choose and call your tools, and Forecall measures how often they get it right.
- Failures reported on your servers: what agents reported on your servers' tools, by error class, tool, model and client.
- Larger allowances for API keys and Failure KB lookups.
Plans and billing compares the plans and their prices.
What static scoring cannot tell
The score reads only the text of each description and schema, so keep these in mind:
- Short tools whose names already say what they do, such as
browser_close, score low even when models use them well. - A high score does not prove that models choose the tool well. It means the text gives them what they need. A low score points to the text worth improving.
- Scores are comparable only within the same version of the scoring rules. Every result shows the version it was scored with.
Try it
Paste the JSON your server returns for tools/list at the linter, or run the CLI
on a file. Get your tools/list shows how to get the JSON.
npx forecall lint tools.json