# Filesystem revalidation race experiment ## Question Can descriptor-relative metadata revalidation ensure that `unlinkat` deletes only the reviewed inode when another process can replace the directory entry? This tests the safety assumption behind [filesystem revalidation](../PRODUCT.md#111-filesystem-revalidation). ## Environment - Date: 2026-09-13 - Kernel: Linux 6.1.176 Android 14, AArch64 - Compiler: Clang 21.1.8 targeting `aarch64-unknown-linux-android24` - Filesystem fixture: a fresh private directory below Termux `$TMPDIR` ## Method A native C probe opened each fixture directory with `O_DIRECTORY | O_NOFOLLOW` and retained its descriptor. It reviewed and revalidated `victim` with `fstatat(directory_fd, "victim", ..., AT_SYMLINK_NOFOLLOW)`, comparing device, inode, type, size, and modification time. It mutated with `unlinkat(directory_fd, "victim", 0)`. The replacement was injected deterministically rather than with timing or sleeps: 1. Review `victim`. 2. Perform the final successful revalidation. 3. Rename `victim` to `saved-reviewed`. 4. Rename an unreviewed regular file to `victim`. 5. Call `unlinkat` on `victim`. Control cases deleted an unchanged victim, replaced the victim before validation, and substituted a symlink after validation. ## Observations ```text control: validation=match reviewed_deleted=yes sentinel_survives=yes replace-before-validation: validation=MISMATCH-BLOCKED replacement_survives=yes replace-after-validation: validation=match unreviewed_replacement_deleted=YES reviewed_survives=yes symlink-after-validation: symlink_deleted=yes target_survives=yes ``` The probe compiled with `-std=c17 -Wall -Wextra -Werror -O2` and exited successfully. ## Conclusion **A final metadata check followed by `unlinkat` does not provide atomic identity-conditional deletion.** A replacement before revalidation is detected, and deleting a substituted symlink does not delete its target. A regular-file replacement in the interval after revalidation is nevertheless deleted. Linux and POSIX expose no `unlink` operation conditional on expected device and inode identity. The specification therefore documents this residual same-permission concurrency limit, requires every observed mismatch to block, and requires disclosure for direct filesystem deletion. This experiment disproves the simple check-then-unlink strategy under concurrent replacement; it does not prove that every more elaborate quarantine protocol is unsafe. ## Quarantine follow-up A second native C probe tested whether moving the selected entry to a mode-`0700` quarantine and verifying it there closes the race. It used `renameat2(..., RENAME_NOREPLACE)` and deterministic operation ordering. A child process with the same UID independently opened the quarantine rather than using inherited descriptors. ```text unlinkat-AT_EMPTY_PATH: result=-1 errno=22 file_survives=yes replacement-before-quarantine: postcheck=MISMATCH unreviewed_moved=YES rollback_result=-1 rollback_errno=17 source_occupied=yes replacement-after-quarantine-check: same_uid_opened_0700=yes unreviewed_deleted=YES reviewed_survives=yes ``` The probe compiled with `-std=c17 -Wall -Wextra -Werror -O2` and exited successfully. The observations mean: - This Android Linux 6.1 kernel rejects `AT_EMPTY_PATH` for `unlinkat` with `EINVAL`. - Quarantine followed by verification prevents deletion of a replacement moved before that verification, but it has already moved an unreviewed object. - Safe rollback is not guaranteed because another object can occupy the original name; `RENAME_NOREPLACE` then fails with `EEXIST`. - Mode `0700` is not an isolation boundary against another process with the same UID. - Replacing the quarantine entry after its final verification recreates the original race and deletes the unreviewed replacement. Quarantine is therefore a useful mitigation only under a narrower concurrency assumption where other processes do not access it. It does not satisfy an unconditional rule that no race may widen the confirmed plan. ## Linux interface research Research on 2026-09-13 covered current Linux man-pages 6.19, current mainline kernel UAPI and VFS sources, and active UAPI feature proposals. - [`unlinkat`](https://man7.org/linux/man-pages/man2/unlink.2.html) accepts a parent descriptor, a pathname, and only `AT_REMOVEDIR`; it accepts no expected identity. - Current mainline [`fs/namei.c`](https://github.com/torvalds/linux/blob/master/fs/namei.c) rejects every `unlinkat` flag except `AT_REMOVEDIR`. - Current [`io_uring` unlink](https://github.com/torvalds/linux/blob/master/io_uring/fs.c) invokes the same pathname operation and adds no identity condition. - [`openat2`](https://github.com/torvalds/linux/blob/master/include/uapi/linux/openat2.h) constrains path resolution only; it does not reserve an entry for later mutation. - [`renameat2`](https://man7.org/linux/man-pages/man2/rename.2.html) can require destination absence or atomically exchange names, but cannot require the source to match a reviewed descriptor. - The UAPI Group lists both [`AT_EMPTY_PATH` for `unlinkat` and `unlinkat2(dir_fd, name, inode_fd)`](https://github.com/uapi-group/kernel-features#unlinking-via-two-file-descriptors) as wanted features, not implemented interfaces. The proposed `unlinkat2` has exactly the missing identity condition. - Kernel developers discussed descriptor-based unlink in [2020](https://kernsec.org/pipermail/linux-security-module-archive/2020-March/019030.html), including unresolved object-versus-directory-entry and locking semantics. - Linux 6.19 [`F_SETDELEG`](https://man7.org/linux/man-pages/man2/F_GETDELEG.2const.html) can delay conflicting rename and unlink operations. It does not give the holder an identity-conditional unlink: the holder must release the delegation, or the kernel eventually breaks it, before mutation proceeds. **No researched interface provides atomic conditional unlink by reviewed identity.** A future `unlinkat2(dir_fd, name, inode_fd)`-style operation could provide the missing primitive, but it is currently only a proposal and is unavailable on the target Android kernel.