dotfiles bugabingas dorkfiles
owned by admin
README Source History Refs Compare Notes Search pi/agent/skills/experimentation/SKILL.md Raw Rendered preview
name: experimentation
description: "Use for throwaway code or artifacts investigating a technical question. Not normal implementation, permanent tests, or artifact-free research."
Experimentation
Temp code = experiment.
Learn → verify → preserve useful evidence → leave the repo clean.
Trigger
Use if task creates/runs throwaway code/data/artifacts:
scratch scripts/files, spikes, probes, PoCs
microbenchmarks, migration dry-runs
one-off parsers/converters/generators
API/dependency/env behavior checks
temp fixtures/logs/output dirs to answer 1 question
Do not use for normal impl/review/tests except exploratory phase.
Protocol
1. Hypothesis
Before code/run:
Q: ...
H: if <condition/change/input> → <observable> because <reason>
Falsifier: false if <specific obs>
Control/baseline: ...
Cmd/input/env: ...
Artifacts: <temp paths>
Repo cleanup: delete <repo paths> unless promoted
External retention: keep/delete <temp, cache, or OS state> because <value>
Unfalsifiable H ⇒ rewrite.
2. Isolation
Default:
outside repo.
tmp="$(mktemp -d "${TMPDIR:-/tmp}/pi-exp.XXXXXX")"
printf '%s\n' "$tmp"
Repo only when needed:
single obvious dir:
.pi/tmp/<slug>/ | tmp/<slug>/
record every ignored/untracked path
git status --short before+after
No scattered scratch files.
3. Provenance
Record if result matters:
exact cmd(s), inputs, sample data
tool/runtime/dependency versions; lockfile state
env/config/flags/params
random seed(s); nondeterminism notes
commit hash if git
timing context:
hw/load/warmup/repeats
Prefer executable protocol over prose:
cmd/test/mise task/script/fixture.
4. Verification ladder
Use ≥2 independent checks when possible.
Predict output before run.
Positive control:
known-good passes.
Negative control:
known-bad fails.
Edge cases:
empty/min/max/weird.
Assertions/invariants fail loud.
Cross-check via 2nd method/tool/manual sample.
Rerun:
deterministic or bounded variance.
Seed randomness or report variance.
Compare baseline/current vs changed.
Explain mismatch; unresolved ⇒ inconclusive.
1 green run = weak evidence.
Log/screenshot without cmd provenance = weak.
5. Promote | summarize | retain | delete
Classify every artifact:
Promote :
durable test/fixture/doc/bench harness/script/repro.
Summarize :
learned fact only → response/issue/commit msg.
Retain :
useful temp/cache/OS artifacts outside the repo.
Record location, purpose, and expected lifetime.
Delete :
repo-local temporary artifacts and external artifacts without future value.
Promotion requirements:
named project location; no local abs paths/debug prints
minimal fixture size
conforms to fmt/lint/test conventions
Cleanup
Before final answer:
Stop processes that are no longer useful.
Delete experiment-created temporary artifacts inside the repository unless promoted.
Preserve useful and safe temp, cache, or OS state outside the repository.
Report retained external state and ask if its footprint or side effects are uncertain.
git status --short if git.
Explain/ask about unknown repository leftovers.
Delete rule:
Only delete explicit paths created by this experiment.
Never delete user data, research cache, repo files, broad globs, or git-clean
without confirmation.
Cleanup means leaving the user's repository clean, not erasing useful external
experimental state.
Final report shape:
Result: supported | falsified | inconclusive
Evidence: cmd(s) + checks
Promoted: <paths + why>
Retained: <external paths/state + why>
Cleaned: <paths>
Remaining: <intentional dirty paths>
Smells
test.py/scratch.js/out.json/log.txt in repo root
hidden generated dir no owner
real data mutated vs copied sample
manual step not captured
bench no warmup/repeats/context
randomness no seed/variance
claim from 1 happy path
broad cleanup:
rm -rf *, git clean -fdx sans consent
---
name: experimentation
description: "Use for throwaway code or artifacts investigating a technical question. Not normal implementation, permanent tests, or artifact-free research."
---
# Experimentation
Temp code = experiment.
Learn → verify → preserve useful evidence → leave the repo clean.
## Trigger
Use if task creates/runs throwaway code/data/artifacts:
- scratch scripts/files, spikes, probes, PoCs
- microbenchmarks, migration dry-runs
- one-off parsers/converters/generators
- API/dependency/env behavior checks
- temp fixtures/logs/output dirs to answer 1 question
Do **not** use for normal impl/review/tests except exploratory phase.
## Protocol
### 1. Hypothesis
Before code/run:
``` text
Q: ...
H: if <condition/change/input> → <observable> because <reason>
Falsifier: false if <specific obs>
Control/baseline: ...
Cmd/input/env: ...
Artifacts: <temp paths>
Repo cleanup: delete <repo paths> unless promoted
External retention: keep/delete <temp, cache, or OS state> because <value>
```
Unfalsifiable H ⇒ rewrite.
### 2. Isolation
Default:
outside repo.
``` bash
tmp="$(mktemp -d "${TMPDIR:-/tmp}/pi-exp.XXXXXX")"
printf '%s\n' "$tmp"
```
Repo only when needed:
- single obvious dir:
`.pi/tmp/<slug>/` | `tmp/<slug>/`
- record every ignored/untracked path
- `git status --short` before+after
No scattered scratch files.
### 3. Provenance
Record if result matters:
- exact cmd(s), inputs, sample data
- tool/runtime/dependency versions; lockfile state
- env/config/flags/params
- random seed(s); nondeterminism notes
- commit hash if git
- timing context:
hw/load/warmup/repeats
Prefer executable protocol over prose:
cmd/test/mise task/script/fixture.
### 4. Verification ladder
Use ≥2 independent checks when possible.
1. Predict output before run.
2. Positive control:
known-good passes.
3. Negative control:
known-bad fails.
4. Edge cases:
empty/min/max/weird.
5. Assertions/invariants fail loud.
6. Cross-check via 2nd method/tool/manual sample.
7. Rerun:
deterministic or bounded variance.
8. Seed randomness or report variance.
9. Compare baseline/current vs changed.
10. Explain mismatch; unresolved ⇒ inconclusive.
1 green run = weak evidence.
Log/screenshot without cmd provenance = weak.
### 5. Promote | summarize | retain | delete
Classify every artifact:
- **Promote** :
durable test/fixture/doc/bench harness/script/repro.
- **Summarize** :
learned fact only → response/issue/commit msg.
- **Retain** :
useful temp/cache/OS artifacts outside the repo.
Record location, purpose, and expected lifetime.
- **Delete** :
repo-local temporary artifacts and external artifacts without future value.
Promotion requirements:
- named project location; no local abs paths/debug prints
- minimal fixture size
- conforms to fmt/lint/test conventions
## Cleanup
Before final answer:
1. Stop processes that are no longer useful.
2. Delete experiment-created temporary artifacts inside the repository unless promoted.
3. Preserve useful and safe temp, cache, or OS state outside the repository.
4. Report retained external state and ask if its footprint or side effects are uncertain.
5. `git status --short` if git.
6. Explain/ask about unknown repository leftovers.
Delete rule:
``` text
Only delete explicit paths created by this experiment.
Never delete user data, research cache, repo files, broad globs, or git-clean
without confirmation.
Cleanup means leaving the user's repository clean, not erasing useful external
experimental state.
```
Final report shape:
``` text
Result: supported | falsified | inconclusive
Evidence: cmd(s) + checks
Promoted: <paths + why>
Retained: <external paths/state + why>
Cleaned: <paths>
Remaining: <intentional dirty paths>
```
## Smells
- `test.py` / `scratch.js` / `out.json` / `log.txt` in repo root
- hidden generated dir no owner
- real data mutated vs copied sample
- manual step not captured
- bench no warmup/repeats/context
- randomness no seed/variance
- claim from 1 happy path
- broad cleanup:
`rm -rf *` , `git clean -fdx` sans consent