AdversaryGate v2.8.0

Merge gate for AI coding agents

No evidence, no merge.

AdversaryGate runs your tests on the code before and after an agent's patch, measures what the patch changed, and answers with one of three decisions. When it could not measure something, the answer is INCONCLUSIVE. A missing measurement never becomes a green check.

Open source, MIT licensed, Python 3.10 or newer. Measures Python/pytest projects today; other languages get INCONCLUSIVE, never MERGE.

decision recorddemo/demo.py
Blockexit 1

The patch breaks a passing test

baseline pass, patch fail · full suite [0, 1]

Inconclusiveexit 2

A clean patch, coverage never measured

diff_coverage_source: none · ratio: null

Mergeexit 0

The same patch, with evidence

coverage 1.0 computed · 5/5 mutants killed · inputs SHA-256'd

The last two rows are the same patch. Same code, same test, same run. Only the evidence differs. This is the output of the demo, which runs as an acceptance test in CI.

“Almost right” is the expensive kind of wrong

In Stack Overflow's 2025 Developer Survey, 66% of developers named “AI solutions that are almost right, but not quite” as a frustration. Source

A pipeline makes it worse when every failure looks the same. Here are four ways a patch can look verified when it is not, and what the gate answers.

  • The test never ran

    An import error or a wrong test ID exits non-zero on both sides, which a simple before/after comparison reads as “nothing changed”.

    Inconclusive the run did not happen, so it proves nothing
  • Coverage was never measured

    A coverage floor fed by a default value or a number someone typed is cleared on every run.

    Inconclusive only a diff plus a coverage report counts
  • The patch rewrote its own test

    The code gets a bug and the assertion, or a helper it imports, is changed to agree with it. Both sides pass.

    Block the baseline's copy of the test runs against the new code
  • The change is in code the tests can't break

    A patch changes C++ or SQL, or deletes a file, and the Python test that passes never touches that code.

    Inconclusive source it cannot judge is never waved through

Four measurements, one decision

No language model runs inside the verifier. The rules are plain code, and they are the same whichever agent wrote the patch.

  1. Runs the same tests on both sides

    Name the tests, or let coverage pick the ones that executed the changed lines. Each runs on the baseline and on the patch, three times per side, on your project's Python. Only a real test failure counts as evidence. Import errors, missing tests, timeouts and runs killed by their limits are unverified.

  2. Measures diff coverage from artefacts

    Computed from your unified diff and the coverage.py JSON report, with test files left out. Both inputs are recorded by SHA-256, so anyone can recompute the number.

  3. Breaks the lines the patch wrote

    Mutation testing on the changed lines only, one mutant per line before any line gets two. The floor reads the lower bound of an 80% confidence interval, so one killed mutant is not enough to call a suite strong.

  4. Re-runs the full suite on both sides

    If the full suite passed on the baseline and fails on the patch, that is a collateral regression, and the answer is BLOCK.

The patch can't grade itself. If it changed a test the baseline already had, or a test helper, or any pytest configuration file, the gate runs the baseline's copy against the new code. A test bent to agree with a bug fails there, and the answer is BLOCK.

Exit codes your CI can branch on

ExitDecisionWhen
0MergeEvery claim verified, diff coverage at least 80%, a mutation score whose 80% lower confidence bound is at least 75% when source changed, and the full suite still passes.
1BlockThe claim's test passed on the baseline and fails on the patch, or the full suite did.
2InconclusiveSomething could not be measured: a harness error, a run killed by its limits, no coverage evidence, too few mutants, a weak or unmeasured suite, or a changed harness file.
3usage errorBad arguments or inputs, such as a coverage number with no evidence behind it.

Evidence of breakage outranks missing evidence: BLOCK beats INCONCLUSIVE, which beats MERGE. Both floors are flags you can change. A passing claim also says how it passed: fixed when the test failed before the patch, no_regression when it passed on both sides.

Add it to a pull request

With base-sha, the Action builds the baseline, the diff and a per-test coverage report itself, then verifies the tests that executed the lines the pull request changed. They run on the Python you set up, with your dependencies.

GitHub Action

name: Verification gate
on: [pull_request]

jobs:
  gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - uses: actions/setup-python@v5
        with:
          python-version: '3.12'
      - run: pip install -r requirements.txt pytest coverage

      - uses: Sanflow10/adversary-gate@v2.8.0
        with:
          base-sha: ${{ github.event.pull_request.base.sha }}

fetch-depth: 0 is required: the base commit has to be in the checkout's history. To name the tests yourself, add claims with one path::test per line.

Command line

pip install "git+https://github.com/Sanflow10/adversary-gate@v2.8.0"

adversary-gate \
  --baseline ./before --patch ./after \
  --python .venv/bin/python \
  --test-path tests/test_auth.py \
  --test-id test_token_expiry \
  --test-id test_refresh \
  --diff changes.diff \
  --coverage-json coverage.json

To see all three decisions first, clone the repository and run python3 demo/demo.py.

Let the agent call it, not configure it

An agent that writes code can call the gate before it reports a task as done. The agent names what to judge. Whoever runs the agent decides how strictly, and against which baseline: whoever picks the baseline picks the oracle.

MCP server for coding agents

pip install "adversary-gate[mcp]==2.8.0"

# Hermes Agent: ~/.hermes/config.yaml
mcp_servers:
  adversary_gate:
    command: "adversary-gate-mcp"
    timeout: 900
    env:
      ADVERSARY_GATE_BASE_REF: "origin/main"
      ADVERSARY_GATE_PYTHON: "/opt/project-venv/bin/python"
      ADVERSARY_GATE_POLICY: "--sandbox bwrap"
    tools:
      include: [verify_repo, gate_policy]

verify_repo judges the repository's working tree, uncommitted changes included. Floors, sandbox, baseline and interpreter come from the environment above, never from the tool call. adversary-gate-mcp --print-hermes-skill prints a skill that tells the agent to fix code rather than tests on BLOCK, and never to report INCONCLUSIVE as done. Works with any MCP client.

Risk triage with Jev

export TYPESAFE_API_KEY=sk-...

adversary-gate ... \
  --diff changes.diff \
  --triage jev

TypeSafe AI's Jev rates the diff's risk in about 100 ms. A confident high raises this run's floors: strength confidence to 0.95, coverage to 0.90. Any other answer, an error or a timeout changes nothing, so a manipulated "low" opens no door. The diff is sent to TypeSafe's API, truncated to 64 KiB, which is why this is opt-in. Jev is in early access; the client follows its published API.

Every bug in the gate is published

A wrong MERGE is the failure this tool exists to prevent, so a bug in it is handled as a security issue. Findings AG-001 to AG-032 are public. Each one says how it was reproduced, what changed, and whether it is still open.

  • AG-032fixed in 2.4.0

    A patch shipped its own .pytest.ini and a plugin that lied only for the exact bytes of its bug, and got MERGE, exit 0. Every test-harness file is now compared tree against tree with the baseline.

  • AG-024fixed in 2.3.0

    A test killed by the memory limit exited 1 and was read as the patch breaking it. It is now a harness death, and every limit is a flag.

  • AG-025fixed in 2.3.0

    The Action replaced the job's Python with its own, so tests ran without the project's dependencies. They now run on yours.

  • AG-021fixed in 2.5.0

    A patch rewrote the claim's test to agree with its own bug and got MERGE, exit 0. Since 2.5.0 the baseline's copy of the test runs against the patch's code, and the attack is BLOCK.

  • AG-012fixed in 2.1.0

    A patch that only changed C++ reached MERGE on a Python test. Source the gate cannot judge now forces INCONCLUSIVE.

  • AG-030fixed in 2.7.0

    "The test failed before and passes now" and "it passed both times" got the same label. They are now fixed and no_regression, and the output says whether a fix was proven.

  • AG-023fixed in 2.7.0

    The mutation score came from at most six mutants with no confidence interval. The floor now reads the lower bound of an 80% Wilson interval: 1 of 1 killed is no longer "strong".

What it does not do

  • Measure languages other than Python. A C++, Rust or SQL change gets INCONCLUSIVE, never a false MERGE. --test-command can run another language's suite through a script that sets up its environment, so a real BLOCK is still possible there.
  • Find tests outside pytest. Automatic discovery reads pytest node IDs from coverage. A suite run through --test-command is named by hand.
  • Contain hostile code. Resource limits and the optional bubblewrap mode contain accidents, not attacks. Run untrusted patches inside a container or a VM.
  • Decide for you. MERGE means the evidence cleared your floors. The merge itself stays with the person who signs the pull request.