Luigit
repositories / dotfiles

dotfiles

bugabingas dorkfiles

owned by admin

pi/agent/skills/experimentation/SKILL.md

Raw
Rendered preview

name: experimentation description: "Use for throwaway code or artifacts investigating a technical question. Not normal implementation, permanent tests, or artifact-free research."

Experimentation

Temp code = experiment. Learn → verify → preserve useful evidence → leave the repo clean.

Trigger

Use if task creates/runs throwaway code/data/artifacts:

  • scratch scripts/files, spikes, probes, PoCs
  • microbenchmarks, migration dry-runs
  • one-off parsers/converters/generators
  • API/dependency/env behavior checks
  • temp fixtures/logs/output dirs to answer 1 question

Do not use for normal impl/review/tests except exploratory phase.

Protocol

1. Hypothesis

Before code/run:

Q: ...
H: if <condition/change/input> → <observable> because <reason>
Falsifier: false if <specific obs>
Control/baseline: ...
Cmd/input/env: ...
Artifacts: <temp paths>
Repo cleanup: delete <repo paths> unless promoted
External retention: keep/delete <temp, cache, or OS state> because <value>

Unfalsifiable H ⇒ rewrite.

2. Isolation

Default: outside repo.

tmp="$(mktemp -d "${TMPDIR:-/tmp}/pi-exp.XXXXXX")"
printf '%s\n' "$tmp"

Repo only when needed:

  • single obvious dir: .pi/tmp/<slug>/ | tmp/<slug>/
  • record every ignored/untracked path
  • git status --short before+after

No scattered scratch files.

3. Provenance

Record if result matters:

  • exact cmd(s), inputs, sample data
  • tool/runtime/dependency versions; lockfile state
  • env/config/flags/params
  • random seed(s); nondeterminism notes
  • commit hash if git
  • timing context: hw/load/warmup/repeats

Prefer executable protocol over prose: cmd/test/mise task/script/fixture.

4. Verification ladder

Use ≥2 independent checks when possible.

  1. Predict output before run.
  2. Positive control: known-good passes.
  3. Negative control: known-bad fails.
  4. Edge cases: empty/min/max/weird.
  5. Assertions/invariants fail loud.
  6. Cross-check via 2nd method/tool/manual sample.
  7. Rerun: deterministic or bounded variance.
  8. Seed randomness or report variance.
  9. Compare baseline/current vs changed.
  10. Explain mismatch; unresolved ⇒ inconclusive.

1 green run = weak evidence. Log/screenshot without cmd provenance = weak.

5. Promote | summarize | retain | delete

Classify every artifact:

  • Promote: durable test/fixture/doc/bench harness/script/repro.
  • Summarize: learned fact only → response/issue/commit msg.
  • Retain: useful temp/cache/OS artifacts outside the repo. Record location, purpose, and expected lifetime.
  • Delete: repo-local temporary artifacts and external artifacts without future value.

Promotion requirements:

  • named project location; no local abs paths/debug prints
  • minimal fixture size
  • conforms to fmt/lint/test conventions

Cleanup

Before final answer:

  1. Stop processes that are no longer useful.
  2. Delete experiment-created temporary artifacts inside the repository unless promoted.
  3. Preserve useful and safe temp, cache, or OS state outside the repository.
  4. Report retained external state and ask if its footprint or side effects are uncertain.
  5. git status --short if git.
  6. Explain/ask about unknown repository leftovers.

Delete rule:

Only delete explicit paths created by this experiment.
Never delete user data, research cache, repo files, broad globs, or git-clean
without confirmation.
Cleanup means leaving the user's repository clean, not erasing useful external
experimental state.

Final report shape:

Result: supported | falsified | inconclusive
Evidence: cmd(s) + checks
Promoted: <paths + why>
Retained: <external paths/state + why>
Cleaned: <paths>
Remaining: <intentional dirty paths>

Smells

  • test.py/scratch.js/out.json/log.txt in repo root
  • hidden generated dir no owner
  • real data mutated vs copied sample
  • manual step not captured
  • bench no warmup/repeats/context
  • randomness no seed/variance
  • claim from 1 happy path
  • broad cleanup: rm -rf *, git clean -fdx sans consent
---
name: experimentation
description: "Use for throwaway code or artifacts investigating a technical question. Not normal implementation, permanent tests, or artifact-free research."
---

# Experimentation

Temp code = experiment.
Learn → verify → preserve useful evidence → leave the repo clean.

## Trigger

Use if task creates/runs throwaway code/data/artifacts:

- scratch scripts/files, spikes, probes, PoCs
- microbenchmarks, migration dry-runs
- one-off parsers/converters/generators
- API/dependency/env behavior checks
- temp fixtures/logs/output dirs to answer 1 question

Do **not** use for normal impl/review/tests except exploratory phase.

## Protocol

### 1. Hypothesis

Before code/run:

```text
Q: ...
H: if <condition/change/input> → <observable> because <reason>
Falsifier: false if <specific obs>
Control/baseline: ...
Cmd/input/env: ...
Artifacts: <temp paths>
Repo cleanup: delete <repo paths> unless promoted
External retention: keep/delete <temp, cache, or OS state> because <value>
```

Unfalsifiable H ⇒ rewrite.

### 2. Isolation

Default:
outside repo.

```bash
tmp="$(mktemp -d "${TMPDIR:-/tmp}/pi-exp.XXXXXX")"
printf '%s\n' "$tmp"
```

Repo only when needed:

- single obvious dir:
  `.pi/tmp/<slug>/` | `tmp/<slug>/`
- record every ignored/untracked path
- `git status --short` before+after

No scattered scratch files.

### 3. Provenance

Record if result matters:

- exact cmd(s), inputs, sample data
- tool/runtime/dependency versions; lockfile state
- env/config/flags/params
- random seed(s); nondeterminism notes
- commit hash if git
- timing context:
  hw/load/warmup/repeats

Prefer executable protocol over prose:
cmd/test/mise task/script/fixture.

### 4. Verification ladder

Use ≥2 independent checks when possible.

1. Predict output before run.
2. Positive control:
   known-good passes.
3. Negative control:
   known-bad fails.
4. Edge cases:
   empty/min/max/weird.
5. Assertions/invariants fail loud.
6. Cross-check via 2nd method/tool/manual sample.
7. Rerun:
   deterministic or bounded variance.
8. Seed randomness or report variance.
9. Compare baseline/current vs changed.
10. Explain mismatch; unresolved ⇒ inconclusive.

1 green run = weak evidence.
Log/screenshot without cmd provenance = weak.

### 5. Promote | summarize | retain | delete

Classify every artifact:

- **Promote**:
  durable test/fixture/doc/bench harness/script/repro.
- **Summarize**:
  learned fact only → response/issue/commit msg.
- **Retain**:
  useful temp/cache/OS artifacts outside the repo.
  Record location, purpose, and expected lifetime.
- **Delete**:
  repo-local temporary artifacts and external artifacts without future value.

Promotion requirements:

- named project location; no local abs paths/debug prints
- minimal fixture size
- conforms to fmt/lint/test conventions

## Cleanup

Before final answer:

1. Stop processes that are no longer useful.
2. Delete experiment-created temporary artifacts inside the repository unless promoted.
3. Preserve useful and safe temp, cache, or OS state outside the repository.
4. Report retained external state and ask if its footprint or side effects are uncertain.
5. `git status --short` if git.
6. Explain/ask about unknown repository leftovers.

Delete rule:

```text
Only delete explicit paths created by this experiment.
Never delete user data, research cache, repo files, broad globs, or git-clean
without confirmation.
Cleanup means leaving the user's repository clean, not erasing useful external
experimental state.
```

Final report shape:

```text
Result: supported | falsified | inconclusive
Evidence: cmd(s) + checks
Promoted: <paths + why>
Retained: <external paths/state + why>
Cleaned: <paths>
Remaining: <intentional dirty paths>
```

## Smells

- `test.py`/`scratch.js`/`out.json`/`log.txt` in repo root
- hidden generated dir no owner
- real data mutated vs copied sample
- manual step not captured
- bench no warmup/repeats/context
- randomness no seed/variance
- claim from 1 happy path
- broad cleanup:
  `rm -rf *`, `git clean -fdx` sans consent