The answer sounded right. The refund happened twice.
An agent can give a convincing final message and still change the wrong object, retry a payment that already committed, ignore a permission boundary, or follow instructions hidden in a support ticket. Grading the last sentence misses all of it.
Verixa grades actions and outcomes. Every tool call runs against a stateful simulated world with real policy, real idempotency and real balances. Faults land exactly where you put them — before or after the commit — and assertions check what the world looks like when the agent is done.
Questions a transcript can't answer
A real product, not a mockup
Captured from a running server with the bundled demo data. See the full tour →
Running in one command
Python 3.11+, no runtime dependencies, no frontend build and no API key for the demo. The demo seeds eleven real simulator runs — ten reference runs and one known regression.
Gate CI on exit codes: 0 passed, 1 error, 3 failed outcome or regression.
git clone https://github.com/zyvorai/verixa && cd verixa
VERIXA_TOKEN=Admin@321 python3 -m verixa serve --demo
# open http://127.0.0.1:8788 — sign in as admin / Admin@321
python3 -m verixa run all --junit artifacts/reference.xml
python3 -m verixa run ticket-injection --agent regression # exits 3
Open, and honest about its limits
Apache-2.0. CI on every push runs the Python suite on 3.11–3.13, a console-to-API harness, real Chromium tests and a container build. Evidence is labelled simulated-tool-state: Verixa proves what your agent did against simulated systems, not that it will behave identically in production. The event hash chain detects edits; it is not a signature.
Bring your own scenarios
New simulated tools, fault types, agent adapters and scenario packs are all welcome. Start with an issue or a pull request.



