--- name: experimentation description: "Use for throwaway code or artifacts investigating a technical question. Not normal implementation, permanent tests, or artifact-free research." --- # Experimentation Temp code = experiment. Learn → verify → preserve useful evidence → leave the repo clean. ## Trigger Use if task creates/runs throwaway code/data/artifacts: - scratch scripts/files, spikes, probes, PoCs - microbenchmarks, migration dry-runs - one-off parsers/converters/generators - API/dependency/env behavior checks - temp fixtures/logs/output dirs to answer 1 question Do **not** use for normal impl/review/tests except exploratory phase. ## Protocol ### 1. Hypothesis Before code/run: ```text Q: ... H: if → because Falsifier: false if Control/baseline: ... Cmd/input/env: ... Artifacts: Repo cleanup: delete unless promoted External retention: keep/delete because ``` Unfalsifiable H ⇒ rewrite. ### 2. Isolation Default: outside repo. ```bash tmp="$(mktemp -d "${TMPDIR:-/tmp}/pi-exp.XXXXXX")" printf '%s\n' "$tmp" ``` Repo only when needed: - single obvious dir: `.pi/tmp//` | `tmp//` - record every ignored/untracked path - `git status --short` before+after No scattered scratch files. ### 3. Provenance Record if result matters: - exact cmd(s), inputs, sample data - tool/runtime/dependency versions; lockfile state - env/config/flags/params - random seed(s); nondeterminism notes - commit hash if git - timing context: hw/load/warmup/repeats Prefer executable protocol over prose: cmd/test/mise task/script/fixture. ### 4. Verification ladder Use ≥2 independent checks when possible. 1. Predict output before run. 2. Positive control: known-good passes. 3. Negative control: known-bad fails. 4. Edge cases: empty/min/max/weird. 5. Assertions/invariants fail loud. 6. Cross-check via 2nd method/tool/manual sample. 7. Rerun: deterministic or bounded variance. 8. Seed randomness or report variance. 9. Compare baseline/current vs changed. 10. Explain mismatch; unresolved ⇒ inconclusive. 1 green run = weak evidence. Log/screenshot without cmd provenance = weak. ### 5. Promote | summarize | retain | delete Classify every artifact: - **Promote**: durable test/fixture/doc/bench harness/script/repro. - **Summarize**: learned fact only → response/issue/commit msg. - **Retain**: useful temp/cache/OS artifacts outside the repo. Record location, purpose, and expected lifetime. - **Delete**: repo-local temporary artifacts and external artifacts without future value. Promotion requirements: - named project location; no local abs paths/debug prints - minimal fixture size - conforms to fmt/lint/test conventions ## Cleanup Before final answer: 1. Stop processes that are no longer useful. 2. Delete experiment-created temporary artifacts inside the repository unless promoted. 3. Preserve useful and safe temp, cache, or OS state outside the repository. 4. Report retained external state and ask if its footprint or side effects are uncertain. 5. `git status --short` if git. 6. Explain/ask about unknown repository leftovers. Delete rule: ```text Only delete explicit paths created by this experiment. Never delete user data, research cache, repo files, broad globs, or git-clean without confirmation. Cleanup means leaving the user's repository clean, not erasing useful external experimental state. ``` Final report shape: ```text Result: supported | falsified | inconclusive Evidence: cmd(s) + checks Promoted: Retained: Cleaned: Remaining: ``` ## Smells - `test.py`/`scratch.js`/`out.json`/`log.txt` in repo root - hidden generated dir no owner - real data mutated vs copied sample - manual step not captured - bench no warmup/repeats/context - randomness no seed/variance - claim from 1 happy path - broad cleanup: `rm -rf *`, `git clean -fdx` sans consent