Luigit
repositories / termux-janitor

termux-janitor

Interactive cleanup assistant for Termux: transparent, safe, confirmed disk reclamation.

owned by admin

spec/experiments/filesystem_revalidation_race.md

Raw
Rendered preview

Filesystem revalidation race experiment

Question

Can descriptor-relative metadata revalidation ensure that unlinkat deletes only the reviewed inode when another process can replace the directory entry?

This tests the safety assumption behind filesystem revalidation.

Environment

  • Date: 2026-09-13
  • Kernel: Linux 6.1.176 Android 14, AArch64
  • Compiler: Clang 21.1.8 targeting aarch64-unknown-linux-android24
  • Filesystem fixture: a fresh private directory below Termux $TMPDIR

Method

A native C probe opened each fixture directory with O_DIRECTORY | O_NOFOLLOW and retained its descriptor. It reviewed and revalidated victim with fstatat(directory_fd, "victim", ..., AT_SYMLINK_NOFOLLOW), comparing device, inode, type, size, and modification time. It mutated with unlinkat(directory_fd, "victim", 0).

The replacement was injected deterministically rather than with timing or sleeps:

  1. Review victim.
  2. Perform the final successful revalidation.
  3. Rename victim to saved-reviewed.
  4. Rename an unreviewed regular file to victim.
  5. Call unlinkat on victim.

Control cases deleted an unchanged victim, replaced the victim before validation, and substituted a symlink after validation.

Observations

control: validation=match reviewed_deleted=yes sentinel_survives=yes
replace-before-validation: validation=MISMATCH-BLOCKED replacement_survives=yes
replace-after-validation: validation=match unreviewed_replacement_deleted=YES reviewed_survives=yes
symlink-after-validation: symlink_deleted=yes target_survives=yes

The probe compiled with -std=c17 -Wall -Wextra -Werror -O2 and exited successfully.

Conclusion

A final metadata check followed by unlinkat does not provide atomic identity-conditional deletion. A replacement before revalidation is detected, and deleting a substituted symlink does not delete its target. A regular-file replacement in the interval after revalidation is nevertheless deleted.

Linux and POSIX expose no unlink operation conditional on expected device and inode identity. The specification therefore documents this residual same-permission concurrency limit, requires every observed mismatch to block, and requires disclosure for direct filesystem deletion. This experiment disproves the simple check-then-unlink strategy under concurrent replacement; it does not prove that every more elaborate quarantine protocol is unsafe.

Quarantine follow-up

A second native C probe tested whether moving the selected entry to a mode-0700 quarantine and verifying it there closes the race. It used renameat2(..., RENAME_NOREPLACE) and deterministic operation ordering. A child process with the same UID independently opened the quarantine rather than using inherited descriptors.

unlinkat-AT_EMPTY_PATH: result=-1 errno=22 file_survives=yes
replacement-before-quarantine: postcheck=MISMATCH unreviewed_moved=YES rollback_result=-1 rollback_errno=17 source_occupied=yes
replacement-after-quarantine-check: same_uid_opened_0700=yes unreviewed_deleted=YES reviewed_survives=yes

The probe compiled with -std=c17 -Wall -Wextra -Werror -O2 and exited successfully. The observations mean:

  • This Android Linux 6.1 kernel rejects AT_EMPTY_PATH for unlinkat with EINVAL.
  • Quarantine followed by verification prevents deletion of a replacement moved before that verification, but it has already moved an unreviewed object.
  • Safe rollback is not guaranteed because another object can occupy the original name; RENAME_NOREPLACE then fails with EEXIST.
  • Mode 0700 is not an isolation boundary against another process with the same UID.
  • Replacing the quarantine entry after its final verification recreates the original race and deletes the unreviewed replacement.

Quarantine is therefore a useful mitigation only under a narrower concurrency assumption where other processes do not access it. It does not satisfy an unconditional rule that no race may widen the confirmed plan.

Linux interface research

Research on 2026-09-13 covered current Linux man-pages 6.19, current mainline kernel UAPI and VFS sources, and active UAPI feature proposals.

  • unlinkat accepts a parent descriptor, a pathname, and only AT_REMOVEDIR; it accepts no expected identity.
  • Current mainline fs/namei.c rejects every unlinkat flag except AT_REMOVEDIR.
  • Current io_uring unlink invokes the same pathname operation and adds no identity condition.
  • openat2 constrains path resolution only; it does not reserve an entry for later mutation.
  • renameat2 can require destination absence or atomically exchange names, but cannot require the source to match a reviewed descriptor.
  • The UAPI Group lists both AT_EMPTY_PATH for unlinkat and unlinkat2(dir_fd, name, inode_fd) as wanted features, not implemented interfaces. The proposed unlinkat2 has exactly the missing identity condition.
  • Kernel developers discussed descriptor-based unlink in 2020, including unresolved object-versus-directory-entry and locking semantics.
  • Linux 6.19 F_SETDELEG can delay conflicting rename and unlink operations. It does not give the holder an identity-conditional unlink: the holder must release the delegation, or the kernel eventually breaks it, before mutation proceeds.

No researched interface provides atomic conditional unlink by reviewed identity. A future unlinkat2(dir_fd, name, inode_fd)-style operation could provide the missing primitive, but it is currently only a proposal and is unavailable on the target Android kernel.

# Filesystem revalidation race experiment

## Question

Can descriptor-relative metadata revalidation ensure that `unlinkat` deletes only the reviewed inode when another process can replace the directory entry?

This tests the safety assumption behind [filesystem revalidation](../PRODUCT.md#111-filesystem-revalidation).

## Environment

- Date: 2026-09-13
- Kernel: Linux 6.1.176 Android 14, AArch64
- Compiler: Clang 21.1.8 targeting `aarch64-unknown-linux-android24`
- Filesystem fixture: a fresh private directory below Termux `$TMPDIR`

## Method

A native C probe opened each fixture directory with `O_DIRECTORY | O_NOFOLLOW` and retained its descriptor.
It reviewed and revalidated `victim` with `fstatat(directory_fd, "victim", ..., AT_SYMLINK_NOFOLLOW)`, comparing device, inode, type, size, and modification time.
It mutated with `unlinkat(directory_fd, "victim", 0)`.

The replacement was injected deterministically rather than with timing or sleeps:

1. Review `victim`.
2. Perform the final successful revalidation.
3. Rename `victim` to `saved-reviewed`.
4. Rename an unreviewed regular file to `victim`.
5. Call `unlinkat` on `victim`.

Control cases deleted an unchanged victim, replaced the victim before validation, and substituted a symlink after validation.

## Observations

```text
control: validation=match reviewed_deleted=yes sentinel_survives=yes
replace-before-validation: validation=MISMATCH-BLOCKED replacement_survives=yes
replace-after-validation: validation=match unreviewed_replacement_deleted=YES reviewed_survives=yes
symlink-after-validation: symlink_deleted=yes target_survives=yes
```

The probe compiled with `-std=c17 -Wall -Wextra -Werror -O2` and exited successfully.

## Conclusion

**A final metadata check followed by `unlinkat` does not provide atomic identity-conditional deletion.**
A replacement before revalidation is detected, and deleting a substituted symlink does not delete its target.
A regular-file replacement in the interval after revalidation is nevertheless deleted.

Linux and POSIX expose no `unlink` operation conditional on expected device and inode identity.
The specification therefore documents this residual same-permission concurrency limit, requires every observed mismatch to block, and requires disclosure for direct filesystem deletion.
This experiment disproves the simple check-then-unlink strategy under concurrent replacement; it does not prove that every more elaborate quarantine protocol is unsafe.

## Quarantine follow-up

A second native C probe tested whether moving the selected entry to a mode-`0700` quarantine and verifying it there closes the race.
It used `renameat2(..., RENAME_NOREPLACE)` and deterministic operation ordering.
A child process with the same UID independently opened the quarantine rather than using inherited descriptors.

```text
unlinkat-AT_EMPTY_PATH: result=-1 errno=22 file_survives=yes
replacement-before-quarantine: postcheck=MISMATCH unreviewed_moved=YES rollback_result=-1 rollback_errno=17 source_occupied=yes
replacement-after-quarantine-check: same_uid_opened_0700=yes unreviewed_deleted=YES reviewed_survives=yes
```

The probe compiled with `-std=c17 -Wall -Wextra -Werror -O2` and exited successfully.
The observations mean:

- This Android Linux 6.1 kernel rejects `AT_EMPTY_PATH` for `unlinkat` with `EINVAL`.
- Quarantine followed by verification prevents deletion of a replacement moved before that verification, but it has already moved an unreviewed object.
- Safe rollback is not guaranteed because another object can occupy the original name; `RENAME_NOREPLACE` then fails with `EEXIST`.
- Mode `0700` is not an isolation boundary against another process with the same UID.
- Replacing the quarantine entry after its final verification recreates the original race and deletes the unreviewed replacement.

Quarantine is therefore a useful mitigation only under a narrower concurrency assumption where other processes do not access it.
It does not satisfy an unconditional rule that no race may widen the confirmed plan.

## Linux interface research

Research on 2026-09-13 covered current Linux man-pages 6.19, current mainline kernel UAPI and VFS sources, and active UAPI feature proposals.

- [`unlinkat`](https://man7.org/linux/man-pages/man2/unlink.2.html) accepts a parent descriptor, a pathname, and only `AT_REMOVEDIR`; it accepts no expected identity.
- Current mainline [`fs/namei.c`](https://github.com/torvalds/linux/blob/master/fs/namei.c) rejects every `unlinkat` flag except `AT_REMOVEDIR`.
- Current [`io_uring` unlink](https://github.com/torvalds/linux/blob/master/io_uring/fs.c) invokes the same pathname operation and adds no identity condition.
- [`openat2`](https://github.com/torvalds/linux/blob/master/include/uapi/linux/openat2.h) constrains path resolution only; it does not reserve an entry for later mutation.
- [`renameat2`](https://man7.org/linux/man-pages/man2/rename.2.html) can require destination absence or atomically exchange names, but cannot require the source to match a reviewed descriptor.
- The UAPI Group lists both [`AT_EMPTY_PATH` for `unlinkat` and `unlinkat2(dir_fd, name, inode_fd)`](https://github.com/uapi-group/kernel-features#unlinking-via-two-file-descriptors) as wanted features, not implemented interfaces. The proposed `unlinkat2` has exactly the missing identity condition.
- Kernel developers discussed descriptor-based unlink in [2020](https://kernsec.org/pipermail/linux-security-module-archive/2020-March/019030.html), including unresolved object-versus-directory-entry and locking semantics.
- Linux 6.19 [`F_SETDELEG`](https://man7.org/linux/man-pages/man2/F_GETDELEG.2const.html) can delay conflicting rename and unlink operations. It does not give the holder an identity-conditional unlink: the holder must release the delegation, or the kernel eventually breaks it, before mutation proceeds.

**No researched interface provides atomic conditional unlink by reviewed identity.**
A future `unlinkat2(dir_fd, name, inode_fd)`-style operation could provide the missing primitive, but it is currently only a proposal and is unavailable on the target Android kernel.