# Score with the CLI

Score a tools/list file on your own machine with forecall lint, get it from your server with forecall dump, connect your AI clients to the Failure KB with forecall setup, and use them in CI.

`forecall lint` scores your tools with the same rules as [the linter](https://forecall.dev/en/lint), on your own
machine. It makes no network requests: no upload, no telemetry, no update check. To get the
tools/list from your server, use [`forecall dump`](#dump). The CLI needs Node.js 22 or later.

```sh
npx forecall lint tools.json
```

## Input

Give one file, or `-` to read standard input:

```sh
cat tools.json | npx forecall lint -
```

The CLI reads the same shapes as the linter, with the same limits:

- An array of tools: `[{ "name": … }, …]`
- An object with a tools array: `{ "tools": [ … ] }`
- A JSON-RPC response: `{ "result": { "tools": [ … ] } }`

Up to 1 MB (1,048,576 bytes) and 200 tools.

Keys must use the MCP wire format (camelCase). Keys in snake_case are reported with where they are
and how to rename them, not converted. [Get your tools/list](https://forecall.dev/en/docs/tools-list) shows how to get
the file from your server.

## Options

| Option | Meaning |
|---|---|
| `--json` | Print the report as JSON, the same report the web app stores |
| `--fail-under <n>` | Exit with 1 when the average score is below `n` (0 to 100) |
| `--lang <en\|ja>` | Language of the report (default: `en`) |
| `-h`, `--help` | Show the help |
| `-v`, `--version` | Show the version and the version of the scoring rules |

The language is never taken from the environment (such as `LANG`), so a report reads the same on
your machine and in CI.

## Output

By default the CLI prints a report for people: the average score, the number of tools and of
confusable pairs, and the version of the scoring rules; then the issues across the server; then
each tool, lowest score first, with its six items and its issues. It uses colors only on a
terminal, and never when `NO_COLOR` is set. [Read your scores](https://forecall.dev/en/docs/reading-scores) explains
the items and the issue codes.

With `--json`, it prints the report the web app stores for a result, with `lintVersion`, the
version of the scoring rules. When the input cannot be scored, it prints `{"error": …}` instead.
Mistakes in the command line itself are printed as text on standard error.

## Get the tools/list: forecall dump

`forecall dump` asks your MCP server for its `tools/list` and prints it as JSON that
`forecall lint` and [the linter](https://forecall.dev/en/lint) accept. It is the only command that talks to the
network: it starts the server you name, or connects to the URL you give.

```sh
# A stdio server: its command after --
npx forecall dump -o tools.json -- npx -y @modelcontextprotocol/server-filesystem .
# A Streamable HTTP server: its URL
npx forecall dump https://example.com/mcp --header "Authorization: Bearer $TOKEN" > tools.json
```

| Option | Meaning |
|---|---|
| `-o`, `--output <file>` | Write the JSON to `file` instead of standard output |
| `--header "Name: value"` | (URL) Send this request header, such as `Authorization`. Repeatable |
| `--env NAME=value` | (command) Set this environment variable for the server, on top of your shell's environment. Repeatable |
| `--cwd <dir>` | (command) Start the server in `dir` |
| `--timeout <seconds>` | Give up after this many seconds (default: 60) |

The JSON is `{"tools": [...]}` with `server` (the server's name and version), `instructions` when
the server has them, and `source`: the command, or the URL without its query, and when it was
taken. It never records headers or environment values. It reads every page of the list. When the
result is over what Forecall scores (1 MiB, 200 tools), it says so on standard error.

## Connect your AI clients: forecall setup

`forecall setup` connects the AI clients on your machine to the [Failure KB](https://forecall.dev/en/kb), the MCP
server at `https://mcp.forecall.dev/mcp` where agents look up known failures of MCP tools and
the workarounds other agents verified. It needs an agent API key from
[the dashboard](https://app.forecall.dev/api-keys): it opens that page and asks you to paste
the key, or takes `--key fc_agent_...`. The key is written into the clients' own configuration
files only.

```sh
npx forecall setup
npx forecall setup --remove
```

For each client it finds, it adds the MCP server `forecall` to the client's configuration,
puts the KB's instructions into the client's global instructions file between
`<!-- forecall:start -->` and `<!-- forecall:end -->`, and, for Claude Code, offers a
`PostToolUse` hook (`forecall hook`) that suggests a `kb_lookup` when an MCP tool of another
server fails. The hook reads the event on standard input and sends nothing anywhere.

| Client | MCP server configuration | Global instructions |
|---|---|---|
| `claude-code` | `~/.claude.json`; the hooks in `~/.claude/settings.json` | `~/.claude/CLAUDE.md` |
| `claude-desktop` | `claude_desktop_config.json`, through `npx -y mcp-remote` | none |
| `cursor` | `~/.cursor/mcp.json` | none: paste the printed block into Settings → Rules |
| `codex` | `~/.codex/config.toml` | `~/.codex/AGENTS.md` |
| `gemini` | `~/.gemini/settings.json` | `~/.gemini/GEMINI.md` |
| `windsurf` | `~/.codeium/windsurf/mcp_config.json` | `~/.codeium/windsurf/memories/global_rules.md` |

With `--sensor` (never by default), `setup` also adds the sensor for Claude Code: a second
`PostToolUse` hook (`forecall hook --sensor`) that Claude Code runs in the background after
each MCP tool call of another server. It redacts the tool's result on your machine with the
KB's own rules (secrets, addresses, paths, identifiers), cuts it to 2,000 characters, and sends
it with the server's and the tool's names, the shape of the arguments (names, types and
lengths, no values) and the client's name to `https://mcp.forecall.dev/sensor`, with the key
in `~/.claude.json`. The server redacts it again. It gives up after 2 seconds, prints nothing
and does not count against your monthly units. Failures it sees that the KB does not know
are reviewed by Forecall before they become records. `setup --remove --sensor` takes out only
the sensor.

Running it again changes nothing; `--remove` takes everything out of every client. Without a
terminal (CI, a pipe) and without `--key`, it prints what to do and exits with 0.

| Option | Meaning |
|---|---|
| `--key <fc_agent_...>` | The agent key; otherwise asked for at the terminal |
| `--client <name>` | Configure this client even if it is not found; repeatable |
| `--hook` / `--no-hook` | Add, or do not add, the Claude Code hook; asked when omitted |
| `--sensor` | Also add the sensor (above); with `--remove`, take out only the sensor |
| `--remove` | Take everything `setup` added out of every client |
| `--dry-run` | Show what would change and change nothing |

## Exit codes

| Code | Meaning |
|---|---|
| 0 | Scored (and, with `--fail-under`, the average is at least `n`); for `dump`, wrote the JSON |
| 1 | The average score is below `--fail-under` |
| 2 | Could not score: a wrong option, a missing or unreadable file, or invalid input; for `dump`, could not get the tools |

Issues alone do not fail the command: short tools whose names already say what they do score low
(see [what static scoring cannot tell](https://forecall.dev/en/docs/getting-started#limits)), so you decide where to
draw the line.

## In CI

To stop a CI job when the average drops, set `--fail-under`. Pin the minor version: every new
version of the scoring rules comes with a new minor version of the CLI, so a pinned job keeps
scoring by the same rules until you update it.

```sh
npx forecall@0.2 lint tools.json --fail-under 60
```

`forecall --version` prints the version of the CLI and of its scoring rules.

To keep each build's scores in the dashboard, send the `tools/list` to the REST API: see
[API keys and the REST API](https://forecall.dev/en/docs/api#save).
