September 8, 2026
Validate Agent Prompts with Sandbox Executions
Learn how to validate agent prompts using sandbox executions in the Run Agents control plane. Test prompts, review logs, and enforce human approval before pro

Validate Agent Prompts with Sandbox Executions
You need a reliable way to test prompts before they drive autonomous work on your agent runtime. Sandbox executions let you run prompts against sample inputs, capture logs, and inspect token usage without touching live systems.
The Agent Command Center provides one control plane to create these tests, store results in a versioned config object, and require human review for any action that could affect the real world.
Key takeaways
- Sandbox tests run inside your agent runtime backend before any production deployment.
- Every config change stays traceable in a single versioned object.
- Outputs that touch external systems always route through the approvals inbox.
- Token counts and cost estimates appear in real time during each test run.
Prepare Your Agent Runtime for Sandbox Testing
Connect the control plane to an isolated runtime instance that mirrors production settings but blocks external calls. Load the same tools and model parameters you use in live schedules. This setup prevents accidental data exposure while still exercising the full execution path your agents will follow.
Limit tool access during tests so no prompt can trigger real changes. Store these boundaries inside the config object so every test run uses identical constraints. Teams that follow the NIST AI Risk Management Framework guidelines often define these boundaries first to reduce downstream compliance gaps.
- Confirm runtime isolation before the first test.
- Load production model parameters into the test config.
- Disable network egress for external APIs.
- Record baseline token usage for comparison.
- Enable detailed execution logs for every step.
- Map each tool to its allowed operations only.
- Verify that schedule triggers remain inactive.
Create Test Cases Inside the Single Config Object
Write test cases that cover edge inputs, role-specific instructions, and potential failure modes. Keep each case tied to the versioned config object so you can replay it after any prompt change. Well-structured cases also surface hidden dependencies between tools and role prompts early.
- List three to five representative inputs per agent role.
- Include prompts that request tool calls.
- Add cases that should trigger approval rules.
- Note expected intermediate outputs for each case.
- Tag cases by risk level for later review.
- Include one case that requests external data retrieval.
- Document the exact model temperature setting used.
Execute Sandbox Runs and Capture Runtime Checks
Trigger sandbox executions directly from the control plane. Watch logs stream in with token counts and intermediate results before the run completes. Real-time visibility helps you spot prompt drift or unexpected tool chaining before the test finishes.
Compare output against your test expectations. Any deviation in behavior or token spend flags the prompt for revision inside the same config object. Following the OWASP Top 10 for Large Language Model Applications helps teams classify which deviations represent prompt injection risks versus simple logic errors.
- Select the agent and test case from the dashboard.
- Start the sandbox execution with a single click.
- Monitor live logs for tool calls and model responses.
- Record final token usage and estimated cost.
- Mark the run pass or fail based on observed behavior.
Compare Sandbox Results Across Config Versions
Use the version history feature to run the same test cases against older and newer prompt versions. This shows exactly how changes affect execution paths and token spend. Side-by-side comparison also reveals whether a prompt edit reduced or increased the number of approval triggers.
| Config Version | Prompt Change | Avg Tokens | Approval Triggered | Notes |
|---|---|---|---|---|
| v1.2 | Added tool limit | 1240 | No | Stable output |
| v1.3 | Expanded role description | 1680 | Yes | Higher spend flagged |
| v1.4 | Added safety clause | 1190 | No | Matches baseline |
Monitor Token Usage and Cost Estimates in Real Time
Track cumulative token consumption across multiple sandbox runs to establish reliable baselines. The control plane surfaces per-execution cost estimates so you can decide whether a prompt revision justifies the added spend before it reaches production schedules.
Teams often set per-role token ceilings inside the config object. When a test run approaches the ceiling, the system pauses and surfaces the intermediate output for quick review.
- Compare token counts against the prior three versions.
- Flag any run that exceeds 150 percent of baseline spend.
- Export cost estimates for budget planning.
- Correlate high token use with specific tool calls.
- Store the final cost figure with the config version.
Route Sensitive Sandbox Outputs to the Approvals Inbox
Any test output that would touch external systems or modify data must land in the approvals inbox. Review, edit, or reject the action before it can proceed to a production schedule. This step keeps human oversight intact even when sandbox results look promising.
Human approval remains mandatory for every real-world action, even after successful sandbox tests. The inbox keeps a record of every decision tied to the config object version.
Apply Runtime Checks Before Moving to Production
Run a final validation pass that includes tool call limits and schedule-specific approval rules. Confirm that failed runs produce clear error traces for later inspection. These checks also verify that the versioned config object contains the exact parameters used in the last successful sandbox pass.
- Verify all tool limits match production settings.
- Confirm approval rules match the target schedule.
- Check that logs include full intermediate outputs.
- Ensure cost estimates stay under defined thresholds.
- Archive the validated config object version.
- Re-run any case that previously triggered an approval.
For more detail on managing these rules across multiple agents, see how a single versioned config object supports controlled execution.
Inspect Failed Sandbox Runs in Execution History
Open the execution history to review every failed test. Trace the exact step where the prompt diverged and note the token count at that point. Patterns in failed runs often point to missing constraints in the role prompt or overly permissive tool definitions.
Use these records to refine the next config version. Failed runs remain linked to the approvals inbox so reviewers can see prior context.
See how to inspect failed agent runs using execution logs for the full workflow.
Decide When Sandbox Results Justify Deployment
Promote a config object to production only after all test cases pass and token usage stays within limits. Keep the human-in-the-loop requirement active for any schedule that performs external actions. A clear decision log inside the control plane supports later audits.
Document the approval decision inside the control plane so future audits can trace the sandbox results that supported deployment.
Conclusion
Sandbox executions give you repeatable evidence that prompts behave as expected before they run on your agent runtime. Validate prompts early, enforce the single config object, and keep every sensitive action behind the approvals inbox.
Next steps
- Create a sandbox runtime instance and connect it to the control plane.
- Build five test cases inside a new versioned config object.
- Run the first set of executions and review token counts.
- Route any external-action output to the approvals inbox.
- Compare results across versions before production promotion.
Start at https://runagents.pro to set up your first sandbox test.
FAQ
How long should a typical sandbox test run take?
Most sandbox executions complete in under two minutes when token limits are set. Longer runs usually indicate an overly broad prompt that needs revision inside the config object.
Can I reuse the same test cases after updating the config object?
Yes. Every test case stays attached to the versioned config object, so you can replay the exact inputs after any prompt, tool, or approval rule change.
What happens if a sandbox run exceeds the token budget?
The control plane stops the execution and records the final token count. You must adjust the prompt or limits in the config object before the next test.
Do sandbox results replace the need for human approval?
No. Sandbox results only validate prompt behavior. Every action that touches the real world still requires explicit approval through the approvals inbox.
How do I compare token usage between two config versions?
Select both versions in the control plane, run the same test cases, and view the side-by-side token and cost estimates that the execution history stores.