05Scope
What the evidence
does and doesn’t say.
- EXAMPLES
The shipped examples and acceptance checks use authored, recorded controls. Bench publishes no live-model scores or comparisons.
- FREEZE
A frozen hash establishes content identity. It is not a prospective registration, a publisher signature or an independent validation.
- REPORTS
Analysis and reports regenerate from preserved artifacts with the pinned runtime. Newly collected stochastic responses need not match earlier ones.
- TESTS
Passing software tests and browser acceptance are bounded software checks, not a methodology audit, peer review or publication.
- EXPORTS
Exported studies are executable packages: inspecting a received study’s results doesn’t run its code, but running or re-grading it does. Run only studies you trust.
- AGENTS
The MCP server is local only; ChatGPT needs a hosted version, which is not part of this release. An agent that runs a study executes that study’s code with your permissions, just like the CLI.
- SOURCE
MIT licensed. Every public release gets a clean public source repository and tag: ToolsEnabled/toolsenabled-bench, tag
v0.3.1.





