Releases 2.4 to 2.8 ยท October 2026
What changed, and what comes next
Five releases in two days, all aimed at one question: can the patch being judged influence its own verdict? Each answer below closed a way it could.
What changed
-
2.8.0Oct 4, 2026
Coding agents can call the gate, but cannot configure it
- MCP server (
adversary-gate-mcp):verify_repojudges a git working tree, uncommitted changes included. A skill for Hermes Agent ships with the package. - The agent names what to judge. The baseline, the interpreter and the floors are set by whoever runs it, because whoever picks the baseline picks the oracle.
- Risk triage with Jev (opt-in): a confident "high risk" raises the floors for that run. No answer can lower them.
- Listed on Glama, with a Docker image tested in CI.
- MCP server (
-
2.7.0Oct 4, 2026
A mutation score is a sample, so it gets a confidence interval
- The strength floor reads the lower bound of an 80% Wilson interval. One killed mutant out of one is no longer "strong"; five out of five is the minimum (AG-023).
- Mutants are spread one per changed line before any line gets two.
- A passing claim says how it passed:
fixedwhen the test failed before the patch,no_regressionwhen it passed both times (AG-030).
-
2.6.0Oct 4, 2026
The oracle covers helpers and the whole suite
--test-supportdeclares helpers whose path doesn't say "test", so they are restored from the baseline too.- The collateral full-suite run also executes the baseline's tests against the new code. A bent test that belongs to no claim is a BLOCK.
-
2.5.0Oct 4, 2026
A rewritten test is judged by the baseline's copy
- When a patch changes a test that already existed, the gate runs the original against the new code. A bug hidden behind a rewritten assertion becomes BLOCK, and an honest test refactor becomes VERIFIED (AG-021).
-
2.4.0Oct 4, 2026
The patch can't configure the pytest that judges it
- A patch could ship a
.pytest.iniand a plugin that lied only for the exact bytes of its bug, and reach MERGE. Every test-harness file is now compared tree against tree with the baseline (AG-032). - One package,
adversary_gate, instead of four top-level modules (AG-027).
- A patch could ship a
Full detail, with how each finding was reproduced: CHANGELOG and findings ledger.
What comes next
In order of how much each one limits the gate today.
-
Biggest gap
Measure languages other than Python
Today a C++, Rust or SQL change is INCONCLUSIVE. Reaching MERGE needs a coverage adapter and, harder, a mutation adapter per language. Rust comes first, through
cargo-mutants. -
Planned
End-to-end runs with a real agent
The MCP server is tested over the real protocol. The next step is a full loop with Hermes Agent: the agent writes a patch, calls the gate, and acts on the answer.
-
Planned
Validate Jev triage against the live API
The client follows TypeSafe's published API and is tested against a local server that speaks it. It has not run against the live service yet.
-
Always open
Find a wrong answer
A wrong MERGE is the failure this tool exists to prevent. If you find one, it becomes the next finding in the public ledger, with a reproduction and a fix.