DocumentationSecurity review

Pilot plan and acceptance criteria

Run a bounded evaluation with observable control outcomes and a clear rollout decision.

Updated 2026-09-22 Read as Markdown
On this page

Define the pilot before installing

Choose one team, one supported integration, and one action to govern in a test environment. Write down the business risk, control owner, technical owner, authorized approvers, permitted data, and criteria for expanding or stopping the pilot. Confirm the features available in the pilot workspace; the Pro trial does not include every enterprise capability.

An example objective is: “Changes to a deployment workflow through the connected coding agent require human review; ordinary repository inspection remains available.” Use a disposable repository and harmless edits to test it.

Phase 1: establish the baseline

Follow Your first governed workflow. Record the agent, platform, client version, hook or endpoint, workspace, environment, and active policy revision. Begin with synthetic data and deliberate shadow mode. Confirm that the integration observes the actions you expect.

Capture representative work that should proceed and a nearby case that should match the proposed restriction. Use these as repeatable test cases throughout the pilot.

Phase 2: test the control

TestExpected observationEvidence to retain
Allowed actionThe harmless operation completes under the intended policyAction result and correlated decision
Denied actionThe targeted operation does not produce its intended side effectDenial, matched rule, and independent target check
Approved HOLDThe action waits, an authorized human approves, and the supported resume flow proceedsCheckpoint, resolution, and final operation result
Denied or expired HOLDThe caller applies the denied or configured timeout outcomeTerminal state, effective decision, and target check
Benign near-matchOrdinary work remains usableInput, observed result, and explanation of the exemption
Policy change and rollbackThe intended revision takes effect and the prior revision can be restoredRevisions, authorized change record, and repeated test
Unavailable service or disconnected hookBehavior matches the documented integration failure postureClient result, failure signal, and residual-risk decision
Credential or membership revocationThe relevant Kastra access is rejected; underlying tool access is reviewed separatelyAuthentication/authorization result and offboarding record
Evidence reviewA reviewer can explain the recorded action and verification scopeExport or decision record, verification result, and coverage notes

Test failure behavior in an isolated environment using your team’s approved method. Run no destructive command merely to demonstrate blocking. For an application dispatcher, additionally test that partial or held model output never triggers a tool.

Phase 3: measure operational fit

Agree on acceptance thresholds before measuring. Useful measures include the proportion of expected test actions observed, false positives among reviewed cases, approval handling time, added end-to-end latency, unresolved failures, and the time a reviewer needs to reconstruct a decision.

Define each denominator and test period. A lower number of recorded events can reflect caching, missing integration coverage, or collection failures. Document those causes instead of treating an incomplete count as reduced risk. Local measurements are evidence for the tested setup, not a service-level commitment.

Make the rollout decision

Summarize the scope tested, results, exceptions, remaining controls, data-use approval, operational owners, and rollback plan. Review the Trust Center and assurance material alongside the pilot results. Expand only to the agent, platform, data, and workflow combinations your team has accepted.

Operational rollout · Discuss a pilot with Kastra