AdversaryGate

Releases 2.4 to 2.8 ยท October 2026

What changed, and what comes next

Five releases in two days, all aimed at one question: can the patch being judged influence its own verdict? Each answer below closed a way it could.

What changed

  1. 2.8.0Oct 4, 2026

    Coding agents can call the gate, but cannot configure it

    • MCP server (adversary-gate-mcp): verify_repo judges a git working tree, uncommitted changes included. A skill for Hermes Agent ships with the package.
    • The agent names what to judge. The baseline, the interpreter and the floors are set by whoever runs it, because whoever picks the baseline picks the oracle.
    • Risk triage with Jev (opt-in): a confident "high risk" raises the floors for that run. No answer can lower them.
    • Listed on Glama, with a Docker image tested in CI.
  2. 2.7.0Oct 4, 2026

    A mutation score is a sample, so it gets a confidence interval

    • The strength floor reads the lower bound of an 80% Wilson interval. One killed mutant out of one is no longer "strong"; five out of five is the minimum (AG-023).
    • Mutants are spread one per changed line before any line gets two.
    • A passing claim says how it passed: fixed when the test failed before the patch, no_regression when it passed both times (AG-030).
  3. 2.6.0Oct 4, 2026

    The oracle covers helpers and the whole suite

    • --test-support declares helpers whose path doesn't say "test", so they are restored from the baseline too.
    • The collateral full-suite run also executes the baseline's tests against the new code. A bent test that belongs to no claim is a BLOCK.
  4. 2.5.0Oct 4, 2026

    A rewritten test is judged by the baseline's copy

    • When a patch changes a test that already existed, the gate runs the original against the new code. A bug hidden behind a rewritten assertion becomes BLOCK, and an honest test refactor becomes VERIFIED (AG-021).
  5. 2.4.0Oct 4, 2026

    The patch can't configure the pytest that judges it

    • A patch could ship a .pytest.ini and a plugin that lied only for the exact bytes of its bug, and reach MERGE. Every test-harness file is now compared tree against tree with the baseline (AG-032).
    • One package, adversary_gate, instead of four top-level modules (AG-027).

Full detail, with how each finding was reproduced: CHANGELOG and findings ledger.

What comes next

In order of how much each one limits the gate today.

  • Biggest gap

    Measure languages other than Python

    Today a C++, Rust or SQL change is INCONCLUSIVE. Reaching MERGE needs a coverage adapter and, harder, a mutation adapter per language. Rust comes first, through cargo-mutants.

  • Planned

    End-to-end runs with a real agent

    The MCP server is tested over the real protocol. The next step is a full loop with Hermes Agent: the agent writes a patch, calls the gate, and acts on the answer.

  • Planned

    Validate Jev triage against the live API

    The client follows TypeSafe's published API and is tested against a local server that speaks it. It has not run against the live service yet.

  • Always open

    Find a wrong answer

    A wrong MERGE is the failure this tool exists to prevent. If you find one, it becomes the next finding in the public ledger, with a reproduction and a fix.