Skip to main content
The OWASP MCP Top 10 names the ten security risks that matter most for MCP. mcpscore checks the part of each risk that is visible on a running server, from outside, without calling a tool. Most of those checks are the Security & Auth category of the score. Two come from elsewhere: one Primitives rule, and the package audit, which scores a published package instead of a running server. This page maps each risk to the rules that cover it and says where the coverage stops.
Every rule line above is a Security & Auth rule. The auth_* rules skip because DeepWiki serves anonymous requests, so there is no authorization flow to grade. A server behind a login grades them; see authenticated servers.

The mapping

A risk is covered in part when mcpscore checks some of what the risk describes. No risk is covered in full: each one also has a part that only code, configuration or the client can show. One Security & Auth rule maps to no OWASP risk. security_malformed_request_handling checks that a malformed request gets the JSON-RPC parse error, which is error hygiene rather than one of the ten risks. The specification sets the same footing for these checks. Its §Security and Trust & Safety says tool descriptions are untrusted unless they come from a trusted server, and that implementers should follow security best practices. The Origin and TLS rules enforce requirements the spec states directly, in Transports §Security Warning and Authorization. Every rule’s citation is in the rules reference.

Where an audit from outside stops

An audit sees what a server publishes and how it answers requests. It never calls a tool, never reads source code and never sees the client. Each gap below follows from one of those limits.
  • MCP01: mcpscore finds a credential the server publishes. It does not see how the server stores tokens, how long they live, or what it writes to its logs.
  • MCP02: mcpscore sees that scopes are advertised. What a granted token can actually do is decided inside the authorization server and the tools.
  • MCP03 and MCP06: the catalog-text rules read for hidden characters and a fixed list of phrasings. A description that misleads in plain words passes. Text returned by a tool is out of reach, because the audit never calls one. A catalog that changes after the audit is out of reach too; running mcpscore in CI catches the change on the next run.
  • MCP04: a package audit reads registry metadata only. It does not download, run or scan the package or its dependencies.
  • MCP05: injection happens when a tool runs with untrusted input. The audit never sends tools/call. Smoke mode calls your own read-only tools with inputs derived from their schemas, not attack payloads.
  • MCP07: mcpscore checks the authorization discovery chain. Whether each tool enforces authorization, and whether the server rejects a token issued for another service, needs credentials and tool calls.
  • MCP08: logging and telemetry live inside the server.
  • MCP09: an inventory of the MCP servers running in an organization is a governance task. mcpscore audits the server you point it at.
  • MCP10: context is shared or kept in the client and in server-side session state, neither of which a request from outside can observe.
For the risks marked Not covered, pair mcpscore with code review, a dependency scanner, and the logging and inventory controls your organization already runs. The testing tools comparison covers which kind of tool answers which question.

How the catalog-text rules decide

The three catalog rules read every string a server publishes: instructions, server info, names, titles, descriptions, schema strings, URIs, MIME types, _meta and icon URLs. A finding names where the string is and what class matched. It never repeats the string, so a report cannot leak a credential or replay an injected instruction.
  • catalog_hidden_unicode flags Unicode tag characters outside a valid emoji tag sequence, bidirectional overrides and isolates, control characters other than tab and line breaks, and runs of two or more zero-width characters, directional marks, soft hyphens or variation selectors. One such character alone passes, because emoji and right-to-left text use them.
  • catalog_no_embedded_secrets matches credentials by their provider’s issued format: AWS access key IDs, GitHub, GitLab and Slack tokens, Stripe live keys, Anthropic, OpenAI and Google API keys, private-key headers, JWTs and long bearer tokens. Documentation samples and placeholder values pass.
  • catalog_prompt_injection_phrasing matches a fixed list of directives to the model: overriding earlier instructions, keeping an action from the user, reassigning the model to a persona, and chat-template control tokens. A phrase in matched quotes, listed with a slash, or introduced as an attempt (“attempts to override the system prompt”) names an attack rather than performing it, and passes. No model judges the text.
  • All three skip a partial audit as insufficient-data, because a server behind a login shows no catalog without credentials.
A finding you intend, such as a game whose instructions give the model a role, can be turned off for your own CI in a mcpscore.toml. The public score on mcpscore.dev keeps every rule.

Security findings in CI

--sarif writes every failed rule as a code scanning alert. Security & Auth rules carry GitHub’s security-severity, so they sort into the Security tab’s critical, high, medium and low bands next to your other security alerts.
In that file, the security_origin_validation finding has level error and security-severity 7.0, which code scanning files as high. The CLI reference has the full mapping. The GitHub Action writes the same file through its sarif-path input.

What’s next