Merge gate for AI coding agents
No evidence, no merge.
AdversaryGate runs your tests on the code before and after an agent's patch, measures what the patch changed, and answers with one of three decisions. When it could not measure something, the answer is INCONCLUSIVE. A missing measurement never becomes a green check.
Open source, MIT licensed, Python 3.10 or newer. Measures Python/pytest projects today; other languages get INCONCLUSIVE, never MERGE.
The patch breaks a passing test
baseline pass, patch fail · full suite [0, 1]
A clean patch, coverage never measured
diff_coverage_source: none · ratio: null
The same patch, with evidence
coverage 1.0 computed · 5/5 mutants killed · inputs SHA-256'd
“Almost right” is the expensive kind of wrong
In Stack Overflow's 2025 Developer Survey, 66% of developers named “AI solutions that are almost right, but not quite” as a frustration. Source
A pipeline makes it worse when every failure looks the same. Here are four ways a patch can look verified when it is not, and what the gate answers.
-
The test never ran
An import error or a wrong test ID exits non-zero on both sides, which a simple before/after comparison reads as “nothing changed”.
Inconclusive the run did not happen, so it proves nothing -
Coverage was never measured
A coverage floor fed by a default value or a number someone typed is cleared on every run.
Inconclusive only a diff plus a coverage report counts -
The patch rewrote its own test
The code gets a bug and the assertion, or a helper it imports, is changed to agree with it. Both sides pass.
Block the baseline's copy of the test runs against the new code -
The change is in code the tests can't break
A patch changes C++ or SQL, or deletes a file, and the Python test that passes never touches that code.
Inconclusive source it cannot judge is never waved through
Four measurements, one decision
No language model runs inside the verifier. The rules are plain code, and they are the same whichever agent wrote the patch.
-
Runs the same tests on both sides
Name the tests, or let coverage pick the ones that executed the changed lines. Each runs on the baseline and on the patch, three times per side, on your project's Python. Only a real test failure counts as evidence. Import errors, missing tests, timeouts and runs killed by their limits are unverified.
-
Measures diff coverage from artefacts
Computed from your unified diff and the coverage.py JSON report, with test files left out. Both inputs are recorded by SHA-256, so anyone can recompute the number.
-
Breaks the lines the patch wrote
Mutation testing on the changed lines only, one mutant per line before any line gets two. The floor reads the lower bound of an 80% confidence interval, so one killed mutant is not enough to call a suite strong.
-
Re-runs the full suite on both sides
If the full suite passed on the baseline and fails on the patch, that is a collateral regression, and the answer is BLOCK.
The patch can't grade itself. If it changed a test the baseline already had, or a test helper, or any pytest configuration file, the gate runs the baseline's copy against the new code. A test bent to agree with a bug fails there, and the answer is BLOCK.
Exit codes your CI can branch on
| Exit | Decision | When |
|---|---|---|
| 0 | Merge | Every claim verified, diff coverage at least 80%, a mutation score whose 80% lower confidence bound is at least 75% when source changed, and the full suite still passes. |
| 1 | Block | The claim's test passed on the baseline and fails on the patch, or the full suite did. |
| 2 | Inconclusive | Something could not be measured: a harness error, a run killed by its limits, no coverage evidence, too few mutants, a weak or unmeasured suite, or a changed harness file. |
| 3 | usage error | Bad arguments or inputs, such as a coverage number with no evidence behind it. |
Evidence of breakage outranks missing evidence: BLOCK beats INCONCLUSIVE, which beats MERGE. Both floors are flags you can change. A passing claim also says how it passed: fixed when the test failed before the patch, no_regression when it passed on both sides.
Add it to a pull request
With base-sha, the Action builds the baseline, the diff and a per-test coverage report itself, then verifies the tests that executed the lines the pull request changed. They run on the Python you set up, with your dependencies.
GitHub Action
name: Verification gate
on: [pull_request]
jobs:
gate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: pip install -r requirements.txt pytest coverage
- uses: Sanflow10/adversary-gate@v2.8.0
with:
base-sha: ${{ github.event.pull_request.base.sha }}
fetch-depth: 0 is required: the base commit has to be in the checkout's history. To name the tests yourself, add claims with one path::test per line.
Command line
pip install "git+https://github.com/Sanflow10/adversary-gate@v2.8.0"
adversary-gate \
--baseline ./before --patch ./after \
--python .venv/bin/python \
--test-path tests/test_auth.py \
--test-id test_token_expiry \
--test-id test_refresh \
--diff changes.diff \
--coverage-json coverage.json
To see all three decisions first, clone the repository and run python3 demo/demo.py.
Let the agent call it, not configure it
An agent that writes code can call the gate before it reports a task as done. The agent names what to judge. Whoever runs the agent decides how strictly, and against which baseline: whoever picks the baseline picks the oracle.
MCP server for coding agents
pip install "adversary-gate[mcp]==2.8.0"
# Hermes Agent: ~/.hermes/config.yaml
mcp_servers:
adversary_gate:
command: "adversary-gate-mcp"
timeout: 900
env:
ADVERSARY_GATE_BASE_REF: "origin/main"
ADVERSARY_GATE_PYTHON: "/opt/project-venv/bin/python"
ADVERSARY_GATE_POLICY: "--sandbox bwrap"
tools:
include: [verify_repo, gate_policy]
verify_repo judges the repository's working tree, uncommitted changes included. Floors, sandbox, baseline and interpreter come from the environment above, never from the tool call. adversary-gate-mcp --print-hermes-skill prints a skill that tells the agent to fix code rather than tests on BLOCK, and never to report INCONCLUSIVE as done. Works with any MCP client.
Risk triage with Jev
export TYPESAFE_API_KEY=sk-...
adversary-gate ... \
--diff changes.diff \
--triage jev
TypeSafe AI's Jev rates the diff's risk in about 100 ms. A confident high raises this run's floors: strength confidence to 0.95, coverage to 0.90. Any other answer, an error or a timeout changes nothing, so a manipulated "low" opens no door. The diff is sent to TypeSafe's API, truncated to 64 KiB, which is why this is opt-in. Jev is in early access; the client follows its published API.
Every bug in the gate is published
A wrong MERGE is the failure this tool exists to prevent, so a bug in it is handled as a security issue. Findings AG-001 to AG-032 are public. Each one says how it was reproduced, what changed, and whether it is still open.
- AG-032fixed in 2.4.0
A patch shipped its own
.pytest.iniand a plugin that lied only for the exact bytes of its bug, and got MERGE, exit 0. Every test-harness file is now compared tree against tree with the baseline. - AG-024fixed in 2.3.0
A test killed by the memory limit exited 1 and was read as the patch breaking it. It is now a harness death, and every limit is a flag.
- AG-025fixed in 2.3.0
The Action replaced the job's Python with its own, so tests ran without the project's dependencies. They now run on yours.
- AG-021fixed in 2.5.0
A patch rewrote the claim's test to agree with its own bug and got MERGE, exit 0. Since 2.5.0 the baseline's copy of the test runs against the patch's code, and the attack is BLOCK.
- AG-012fixed in 2.1.0
A patch that only changed C++ reached MERGE on a Python test. Source the gate cannot judge now forces INCONCLUSIVE.
- AG-030fixed in 2.7.0
"The test failed before and passes now" and "it passed both times" got the same label. They are now
fixedandno_regression, and the output says whether a fix was proven. - AG-023fixed in 2.7.0
The mutation score came from at most six mutants with no confidence interval. The floor now reads the lower bound of an 80% Wilson interval: 1 of 1 killed is no longer "strong".
What it does not do
- Measure languages other than Python. A C++, Rust or SQL change gets INCONCLUSIVE, never a false MERGE.
--test-commandcan run another language's suite through a script that sets up its environment, so a real BLOCK is still possible there. - Find tests outside pytest. Automatic discovery reads pytest node IDs from coverage. A suite run through
--test-commandis named by hand. - Contain hostile code. Resource limits and the optional bubblewrap mode contain accidents, not attacks. Run untrusted patches inside a container or a VM.
- Decide for you. MERGE means the evidence cleared your floors. The merge itself stays with the person who signs the pull request.