We launched ToolBench because we repeatedly heard teams building agents ask if an MCP server was good enough for production. Our goal was to give builders and buyers an objective score to help answer that question, and give tool builders a clear list of what to fix.
Six months after our initial launch we finished a full rescan. ToolBench now scores more than 90,000 open-source MCP servers exposing more than 630,000 tools. That’s over twice the servers and nearly three times the tools we covered at launch, both numbers that indicate just how rapidly the ecosystem is growing.
While that overall growth is encouraging, ToolBench once again shows us just how much work there is still to be done.
MCP settled into enterprise work
In just six months MCP has solidified as the protocol enterprises build agents on, and the July 28 specification was a major validation. The new spec makes MCP stateless, with no handshake and no session ID, so any request can reach any server instance. It formalizes Tasks and MCP Apps as official extensions and hardens authorization. AWS, Cloudflare, and Microsoft all endorsed the release.
The servers reflect the shift. Among those we could classify by industry, about 30% sit in software, infrastructure, data, or security, and another 20% serve business functions like finance, sales, support, and HR. These are the systems organizations run on.
We’re also starting to see the new spec in the wild. Stateless handling appears on 5.8% of servers, and those servers average about eight points higher. Just over 300 servers have adopted multi round-trip requests.

Adoption is early, and spec version alone doesn’t predict quality, but the builders moving tend to be the ones doing other things right.
More servers, same unfinished work
Scale hasn’t solved quality. Only 260 of the open-source servers we scored (0.28%) earn an A, and none earn an A+. We rebuilt the rubric for this rescan around the July 28 spec, so grades aren’t directly comparable to earlier scans. Raising grades across the ecosystem will take real work, and the data shows that work has shifted in six months.
Builders learned to name tools before they learned to explain them
When we first launched ToolBench in March, missing descriptions topped the issue list. Today the average tool name scores 79, while descriptions and input schemas both average 57. Builders have learned naming conventions. They haven’t yet learned to tell an agent what a tool changes, what it returns, and what inputs it accepts. Input-constraint problems show up on 53% of servers.

The largest gap sits downstream. Seven in 10 servers give agents no useful guidance when a call fails. Agents find the right tool, then fall back on retries and guesswork when something breaks.
Maintained servers are pulling ahead
Some of the ecosystem is getting better. Of the 35,741 servers we scored in both scans, 10,264 have shipped new code since our first scan. Their tool definition scores rose nearly three times as much as those of untouched servers. About 2,800 of these maintained servers improved their definitions.
Most servers, though, haven’t moved. 71% of the servers from our first scan haven’t shipped a line of code since. Much of the ecosystem is set-and-forget, and supportability averages just 36 out of 100.
The riskiest tools get the least documentation
Documentation gets weaker exactly where it matters most. Average description scores fall from 57.8 for read-only tools to 54.0 for destructive tools and 52.9 for irreversible ones, even within the same server. Across nearly 17,000 servers that ship a destructive or irreversible tool, only 14% declare safety annotations, and just 101 use elicitation, the protocol’s built-in way to ask the user before acting.

Mislabeling is rare. Only 57 of more than 29,000 destructive or irreversible tools are named like reads. But about half use neutral verbs such as run, clear, or reset, and most never warn about consequences in the description.
That means an agent has no way to know a call will delete data until it does.
Popularity and polish don’t close the safety gap
Repository maturity helps the score. A detected license is associated with about eight extra points and a documentation site with about seven, while more than 10 GitHub stars adds about two. The strongest signal is interface design itself: servers with tool annotations score nearly 11 points higher.
Maturity doesn’t fix recovery or safety, though. Among servers with recorded findings, those with both a license and documentation still get flagged for missing recovery guidance about 92% of the time, nearly the same rate as servers with neither. Servers that return structured errors fare no better, because telling an agent a call failed doesn’t tell it what to do next. Credential-handling problems hold steady at 5% to 7% across every star band.
Better servers face harder problems
As grades rise, the issues change. Description problems drop from 89% of F-grade servers to 61% of A-grade servers, while pagination issues climb from 27% to 46% and missing safety annotations from 10% to 31%. Weaker servers struggle to explain a tool. Stronger servers struggle to operate one safely at scale.
Size matters too. Quality peaks at 16 to 31 tools per server and drops about six points past 32.

What to fix next
If you’re building MCP servers and agent-ready tools, these fixes have the clearest payoff in the data.
- Tell the agent what the tool does, what it changes, and what it returns. Say when to use it and when not to; descriptions that do score 3 to 9 points higher at the same length, and only about 3% of tools do it today.
- Constrain inputs with enums, formats, and ranges.
- Return errors that tell the agent whether to retry, fix its input, or stop.
- Annotate destructive tools and ask the user before irreversible actions.
- Keep credentials out of tool parameters, where they end up in prompt history and logs.
- Grow a toolkit only as fast as its definitions stay coherent.
See where your servers stand at ToolBench, and read our full methodology here.