python-code-validator
Validate, repair and sandbox-run AI-generated Python.
Documentation
Python Code Validator
**An MCP server that validates, repairs and runs Python against the examples it
is supposed to satisfy — `validate_python`, `repair_python` and `execute_python`
over HTTP at `https://api.statemind.ai/mcp`, with a free key and no account.**
A hosted service that proves AI-generated Python does what you asked. State the
intent — assertions or doctest lines — and the code is run against it inside a
container with no network and a read-only filesystem; a fix comes back only when
every example passes. On the QuixBugs defects that is 41% repaired and 77%
refused as not doing what they say, with no false alarms on the corrected
programs — where `ruff` and `mypy` flag the defect in none of them
(the numbers).
The checks that need no intent come with it: syntax and lint diagnostics, an AST
security policy that also catches calls hidden behind dynamic imports and runtime
attribute lookups, a bandit pass, a credential scan and deterministic repair —
one verdict with a score. Asking the same question twice inside ten minutes is
answered from the first answer and costs nothing (`x-msvc-repeat: 1`).
This repository holds the client side: the MCP configuration, the CI script and
the pre-commit hook. The service itself runs at `https://api.statemind.ai`, so
there is nothing to install or host.
A key, without an account
curl -s -X POST https://api.statemind.ai/v1/keys
# {"api_key": "msvc_free_…", "tier": "free", "calls_per_day": 25, "modes": ["static"]}25 static checks a day, metered per UTC day, and a few keys per address: enough
to try it and to run it over a small project, not a supply. Every answer carries
the state of the allowance (`x-quota-remaining`, `x-quota-reset`), so a client
can back off before it is cut off.
MCP
Registered in the official MCP registry as
`ai.statemind/python-code-validator`, a name verified against the domain that
serves it rather than a GitHub account. Any MCP client adds it with one
block:
{
"mcpServers": {
"python-code-validator": {
"type": "http",
"url": "https://api.statemind.ai/mcp",
"headers": { "Authorization": "Bearer msvc_free_…" }
}
}
}- Claude Code: `claude mcp add --transport http python-code-validator https://api.statemind.ai/mcp --header "Authorization: Bearer msvc_free_…"`
- Cursor: `~/.cursor/mcp.json`, same block.
- VS Code / Copilot: `.vscode/mcp.json` under `"servers"`.
A client that only launches a command uses the stdio bridge in this repository
instead, which forwards the same tool over HTTPS:
{
"mcpServers": {
"python-code-validator": {
"command": "python3",
"args": ["/path/to/python-code-validator/mcp_stdio.py"]
}
}
}Or as a container, which the `Dockerfile` here builds:
docker build -t python-code-validator .
docker run -i --rm -e VALIDATOR_API_KEY python-code-validatorGemini CLI installs the same bridge as an extension, with the instruction file
that makes it get used:
gemini extensions install jkanselaar/python-code-validatorThree tools, named after what they do to the code:
| tool | runs the code | key |
|---|---|---|
| `validate_python` | no | free |
| `repair_python` — also returns `fixed_code` | no | paid |
| `execute_python` — also runs it in a sandbox | yes | paid |
The old single `python_code_validator` tool, with its `mode` argument, still
answers for clients that already configured it, but is no longer listed.
Saying what the code was supposed to do
Every check above passes on a function that computes the wrong answer. The one
thing that catches it is the intent, and the agent that asked for the code is
the only one who has it — so pass it along:
{"code": "def bitcount(n): …", "mode": "execute",
"options": {"examples": "assert bitcount(127) == 7"}}Doctest lines (`>>> bitcount(127)` then `7`) work the same way, as do `>>>`
examples already written in the source. `execute_python` runs them in the
sandbox: one that does not hold is a `python:example-mismatch` error, and the
repair search returns a fix only when every example passes. On the QuixBugs
defect set — real bugs, hidden test inputs deciding correctness — that repairs
41% and refuses 77% as not doing what they say, with no false alarms on the
corrected programs.
Repeating a call costs nothing: the same key asking the same question — same
mode, same code, same examples — is answered from the answer it already got,
marked `x-msvc-repeat: 1`, so an agent that checks its work at every step is not
billed for verdicts that cannot have changed.
Claude Code plugin
An instruction can be ignored; a hook cannot. The plugin checks every Python
file Claude Code writes or edits, in the turn it was written, and hands the
errors back to the model instead of to you:
/plugin marketplace add jkanselaar/python-code-validator
/plugin install python-code-validator@statemindNothing to configure: it mints and keeps its own free key on first use. A file
that comes back accepted is silent, a rejected one stops the turn with the
offending lines named, and an identical file is not asked about twice. It never
ends a session over its own trouble — an unreachable service or a spent
allowance lets the turn continue, and the allowance says how to raise it.
Set `VALIDATOR_API_KEY` to use a paid key instead of the free tier, and
`VALIDATOR_URL` to point at your own deployment. The plugin also carries the
`validate-python` skill, for the part a hook cannot do: stating the intent as
examples and running the code against them.
Cursor hook
The same script, wired to Cursor's `postToolUse`, where the verdict comes back
as context on the conversation instead of as an exit code:
mkdir -p .cursor/hooks
base=https://raw.githubusercontent.com/jkanselaar/python-code-validator/main
curl -sf $base/plugin/hooks/validate_written.py -o .cursor/hooks/validate_written.py
curl -sf $base/cursor/hooks.json -o .cursor/hooks.jsonProject hooks run from the project root, which is why the command in
`cursor/hooks.json` is a path relative to it. For a hook
that applies to every project instead, put the script in `~/.cursor/hooks/` and
the same block in `~/.cursor/hooks.json` with the command
`python3 ./hooks/validate_written.py --cursor`.
Making the agent use it
Configuring the server is not what gets it called: the instruction file is.
`AGENTS.md` in this repository is that text, written to be dropped
into any project under whichever name the client reads:
mkdir -p .github
curl -sf https://raw.githubusercontent.com/jkanselaar/python-code-validator/main/AGENTS.md \
| tee AGENTS.md CLAUDE.md GEMINI.md .github/copilot-instructions.md >/dev/nullCursor reads rules with front matter instead, so that one is a separate file —
copy `.cursor/rules/python-code-validator.mdc`
into `.cursor/rules/` of the project.
The short version, if you would rather add a line to instructions you already
have:
> Write what the code should do as `assert` examples before writing the code,
> and pass them in `options.examples`. Call `validate_python` after every edit
> and `execute_python` once a function is finished, not again until what it
> does has changed. When a call returns `fixed_code`, take it — the service ran
> it against your examples. Do not present code that came back `valid: false`.
CI
The service hands out the client, so a workflow needs no checkout of this
repository and no secret:
- run: |
curl -sf https://api.statemind.ai/v1/client -o validate.py
python3 validate.py --changed-against "origin/${{ github.base_ref }}"Or as an action, from the Marketplace:
permissions:
contents: read
pull-requests: write # so the run can comment its result on the pull request
steps:
- uses: jkanselaar/python-code-validator@v1.22.0
with:
api-key: ${{ secrets.VALIDATOR_API_KEY }} # optional; free tier without itThe changed Python is validated and offending lines are annotated on the diff,
failing the job on syntax errors and unsafe patterns. Files the service refuses
outright (over its 200 kB limit) are skipped with a warning rather than failing
the run.
The run also leaves one comment on the pull request, edited in place on later
pushes rather than repeated: what was accepted, what was repaired and how much of
the day's allowance is left. Without `pull-requests: write` nothing is written
and the job is unaffected; `comment: "false"` turns it off.
On the free tier the action keeps its key in the workflow cache, one per
repository per day, so the allowance belongs to the repository rather than to the
run. With `api-key` set the cache is skipped.
Pre-commit
repos:
- repo: https://github.com/jkanselaar/python-code-validator
rev: v1.22.0
hooks:
- id: python-code-validatorThe client itself
`validate.py` is standard library only, so it also works as `python
validate.py file.py` in a Makefile, a git hook or a container:
$ python3 validate.py service.py
::error file=service.py,line=88,title=SyntaxError::invalid syntax
FAIL service.py score=0.66
0/1 files accepted`VALIDATOR_API_KEY` is used when set; otherwise the client mints a free key —
keeping it in `VALIDATOR_KEY_FILE` when that names a path, which is how a series
of runs shares one allowance. `VALIDATOR_URL` points it at another deployment.
`VALIDATOR_SOURCE` names the caller, which is only ever counted: a run inside a
workflow says `github-action` by itself.
The badge
A repository whose Python is checked on every pull request can say so:
[](https://api.statemind.ai/?src=badge)HTTP
curl -s https://api.statemind.ai/v1/validate \
-H "Authorization: Bearer $VALIDATOR_API_KEY" \
-H 'content-type: application/json' \
-d '{"code": "def f(:\n pass\n", "mode": "static"}'`mode` is `static`, `repair` or `execute`; `repair` and `execute` need a
configured key. Submitted code is not logged.
A refused call says what to do about it, so a caller with no operator to ask can
resolve it itself:
{"error": "payment_required",
"remedy": {"action": "upgrade_key", "hint": "A free key covers static only. …"}}Paying for calls
A free key covers 25 static checks a day, and one address gets a few keys a day,
so the allowance is a trial rather than a supply. Beyond it a key carries
credits: a static check costs 1, a repair 3 and a sandboxed run 10, and an
identical call repeated within ten minutes is answered from the first one for
free.
Credits are bought with a card, without an invoice or anyone to ask:
curl -s -X POST https://api.statemind.ai/v1/keys/checkout \
-H 'content-type: application/json' \
-d '{"api_key": "'"$VALIDATOR_API_KEY"'", "credits": 500}'That answers with a Stripe Checkout page; the credits are on the key seconds
after the card clears (500 credits is €10). An agent with a Gnosis wallet can
instead pay in xDAI without a browser — `GET /v1/pricing` states both routes.
Examples
`examples/` holds three files and the client to send them with: one
that passes every check and still returns the wrong number, one the security
policy refuses, and one that comes back accepted from the sandbox.
Licence
MIT.
Frequently asked questions
What is python-code-validator?
python-code-validator is Validate, repair and sandbox-run AI-generated Python.
How do I install python-code-validator?
Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.
Is python-code-validator open source?
Yes — it is hosted on GitHub at https://github.com/jkanselaar/python-code-validator.
Related MCP tools
Open-source coding agent memory. Records issues, attempts, fixes and decisions, then warns your agent before it repeats an approach that already failed. Native MCP server for Claude Code, Cursor, Antigravity and Codex. 100% local, no cloud, no telemetry. MIT.
AI-powered OSINT agent with interactive REPL, MCP server, and CLI. 19 tools. Works with Claude, GPT-4, or local models. For authorized security research only.
Build effective agents using Model Context Protocol and simple workflow patterns Python-based implementation. Trusted by 7600+ developers.
Fast and Accurate Code Search for Agents. Uses 99% fewer tokens than grep+read
An AI Gateway, registry, and proxy that sits in front of any MCP, A2A, or REST/gRPC APIs, exposing a unified endpoint with centralized discovery, guardrails and management. Optimizes Agent & Tool calling, and supports plugins.
Cut AI token costs 95%+ on code exploration. The leading MCP server for precise, symbol-level GitHub code retrieval via tree-sitter AST. Works with Claude Code, Cursor & any MCP client. 313B+ tokens saved.
Run your own MCP server? See who uses it and what to fix.
Measure it with TrackMCP