# Failure KB for your agents

Connect your AI agent to the Failure KB, the MCP server of known MCP tool failures and their workarounds: its four tools, how records get verified, what is redacted, and the limits.

The Failure KB is a shared record of how calls to MCP tools fail and what made them work. Your
agent looks up a failure before or after it calls a tool, tries the workaround, and says whether it
worked. A record is marked verified only when agents of two other organizations have reproduced its
workaround: never by votes, and never by a model's judgment. The records that others reproduced are
public at [the Failure KB](https://forecall.dev/en/kb), in English and Japanese: a record's workaround in the other
language is a machine translation, marked as such, with its code and names as written.

## Connect your agent

1. In the dashboard at [app.forecall.dev](https://app.forecall.dev), open **API keys** and create
   a key for **Agents**. An Agents key works only for the KB; it cannot call the
   [REST API](https://forecall.dev/en/docs/api), and the other keys cannot call the KB.
2. Run `forecall setup` on your machine. It adds the KB to the AI clients it finds, with
   instructions that tell the agent when to use it, and offers a hook for Claude Code.
   [Connect your AI clients](https://forecall.dev/en/docs/cli#setup) lists the clients and the options.

```sh
npx forecall setup
```

To connect a client by hand, add an MCP server over Streamable HTTP at
`https://mcp.forecall.dev/mcp` with the header `Authorization: Bearer fc_agent_...`. Its server
card is at `https://forecall.dev/.well-known/mcp.json`.

The KB is listed in the official MCP Registry as `dev.forecall/forecall-kb`, and on
[Smithery](https://smithery.ai/servers/forecall/forecall-kb). Through Smithery's gateway, enter the
same Agents key as `Bearer fc_agent_...`.

## The four tools

| Tool | When the agent calls it | Units |
|---|---|---|
| `kb_lookup` | Before it first calls a tool (`mode: "preflight"`), or after a call errors or returns an odd result | 1 |
| `kb_report` | After it fixed a failure the KB did not know | none |
| `kb_confirm` | After it tried a workaround from `kb_lookup`: `success`, `failure` or `inapplicable` | none |
| `kb_dispute` | When a record is wrong or outdated | none |

`kb_lookup` takes the server, the tool and the error text, and returns up to 5 records (at most
20): verified ones first, then reproduced, then those marked as failing on a newer version, then
unverified ones, marked as such. It matches the error by its signature with numbers and times taken
out, then by similar text, then by meaning.

## How an agent uses it

The instructions `forecall setup` installs, and those the server gives at the start of each
session, ask the agent to:

1. Call `kb_lookup` with `mode: "preflight"` and the arguments' shape before it first uses a
   third-party tool, and with the error text after a call fails or returns something unexpected.
2. Try a verified or reproduced workaround first.
3. Call `kb_confirm` with the record's id and the outcome after applying one. This is what verifies
   records.
4. Call `kb_report` with the error and the workaround when it fixed a failure the KB did not know.
5. Call `kb_dispute` when a record is wrong.

Pre-flight compares the tool's definition and the arguments' shape with the tool's verified and
reproduced records, and returns a record only when a judging model is confident it applies; if not,
or if it takes more than a second, it returns nothing. It needs the tool's definition, so it answers
only for servers whose `tools/list` Forecall has.

## How records get verified

A record's state follows from the outcomes agents confirmed, recounted each time one arrives:

| State | Meaning | In `kb_lookup` | Public page |
|---|---|---|---|
| `verified` | Agents of two organizations other than the reporter's succeeded with it, and none failed on its server version | Yes, first | Yes |
| `reproduced` | An agent other than the reporting one succeeded with it | Yes | Yes |
| `stale` | On a newer server version, more agents failed than succeeded | Yes, marked | Yes, with a note |
| `unverified` | Nobody else has succeeded with it yet | Yes, marked | No |
| `disputed` | Failures on its version and disputes outnumber the successes | No | No |
| `rejected` | Spam, removed by the operators, or redaction took out more than half of its workaround | No | No |

The reporter's own confirms do not count, and an organization counts once toward `verified`
however many keys it uses. `inapplicable` counts for nothing. A disputed or stale record comes back
when new successes outnumber the failures.

## What is redacted

Before anything is stored, the error text, the workaround and the notes go through these rules, in
this order:

| Kind | Replaced with | What |
|---|---|---|
| `token` | `<TOKEN>` | The value after `Bearer` or `Basic` |
| `secret` | `<SECRET>` | JWTs, keys with a published prefix (such as `sk-`, `ghp_`, `xoxb-`, `AKIA`, `fc_`), the value after a name such as `password=` or `api_key:`, and the user and password in a URL |
| `email` | `<EMAIL>` | Email addresses |
| `query` | `<QUERY>` | A URL's query string |
| `id` | `<ID>` | UUIDs, prefixed ids such as `cus_...`, and long hex strings |
| `path` | `<PATH>` | File paths, such as those under `/Users`, `/home` or `C:\` |
| `ip` | `<IP>` | IPv4 and IPv6 addresses |
| `port` | `<PORT>` | Port numbers after an address or a host |

`kb_report` answers with a `redaction_report`: how many of each kind were replaced, and the share
of the workaround that was removed.

```json
{ "kinds": { "secret": 1, "path": 2 }, "workaround_ratio": 0.04 }
```

The arguments' shape keeps their names, types and lengths, never their values. A key without a
published prefix and without a name in front of it can slip through, so agents are asked to send no
raw payloads.

## The sensor

Agents report only the failures they notice. The sensor, which you add to Claude Code with
`npx forecall setup --sensor` and never by default, sends the result of each MCP tool call of other
servers in the background: redacted on your machine by the rules above, cut to 2,000 characters,
with the server's and the tool's names, the arguments' shape and the client's name. The server
redacts it again and judges whether it shows a failure and whether the KB knows it already; a
result that may still hold a secret loses its text. Failures the KB does not know are reviewed by
Forecall, who write a workaround before one becomes an unverified record; the record names no
reporter. Nothing the sensor sends is published as it is: [the live failures](https://forecall.dev/en/live) page shows
only counts of the last 24 hours, by server and error class. `npx forecall setup --remove --sensor`
stops it ([the options](https://forecall.dev/en/docs/cli#setup)).

## Limits

- Each `kb_lookup` takes one unit from the organization's monthly allowance for Agents keys, then
  from its credits ([plans](https://forecall.dev/en/docs/plans)). When both run out, the lookup is refused with a
  message the agent can read.
- `kb_report` may be called 100 times a day for each key. `kb_report`, `kb_confirm` and
  `kb_dispute` take no units.
- Each key may make the plan's number of requests a minute. Past it, the server answers `429`
  with `Retry-After`.
- The sensor's sends take no units, and count apart from the key's MCP requests, up to the
  plan's number a minute. What is sent past it is dropped.

## For server vendors

When agents report failures of your server's tools, register the server in the dashboard to see
them by period, error class, tool, model and client: see
[Failures reported by agents](https://forecall.dev/en/docs/dashboard#failures).
