Luigit
repositories / bugabinga.net

bugabinga.net

personal infrastructure for bugabinga!

owned by admin

.system/research/BB-RESEARCH-F1ZMINXT-luigit-search-indexing-freshness-caching-and-gar/index.md

Raw
Rendered preview

id: BB-RESEARCH-F1ZMINXT type: research title: Luigit search indexing freshness caching and garbage collection

Luigit search indexing freshness caching and garbage collection

Scope

Determine Luigit's searchable representation, update model, cache boundaries, and derived-state reclamation. Preserve current public reachability, seconds-scale update visibility, bounded resources, and browsing during search failure or rebuild.

Facts

Search acceleration is not authorization

Git objects are immutable and content-addressed. Their public status is mutable because branches, tags, notes refs, hosting visibility, and advertisement policy change. Force-pushed and deleted-ref objects may remain in packs, reflogs, and old indexes after losing public reachability.

Index membership, object existence, hidden refs, and reflogs cannot authorize output. This applies to snippets, result counts, facets, suggestions, pagination, and identifiers, not only result links.

Git ref updates are individually atomic. A concurrent reader may observe only part of a multi-ref transaction. A read-only consumer cannot know an instantaneous repository state that changes concurrently; it can authorize against a named successfully observed ref snapshot.

Sources: Git data model, Git repository layout, Git revisions, Git update-ref transactions.

Advertised refs need an explicit snapshot

Git protocol v2 ls-refs can report refs, symbolic refs, and peeled tags while applying hidden-ref policy. The repository's complete ref namespace is not necessarily its public advertisement.

A Luigit visibility snapshot must contain the public repository metadata generation, sorted advertised ref names and targets, peeled tag targets, visible HEAD symbolic target, selected notes refs, object format, and policy version. Its digest identifies the authorization observation used by a request and the indexed coverage built from it.

Code reachability and notes-state reachability are separate. A visible notes ref authorizes its own note history, not the annotated target. A note result remains visible only when its note state is permitted and its target is reachable from current public code roots.

Sources: Git protocol v2, Git hidden-ref configuration, Git notes, Git garbage collection.

Indexed entities and public occurrences are different data

Blob text, commit messages, and note blobs are immutable searchable entities identified by repository, object format, object ID, and kind. Repository names, refs, paths, branch or tag snapshots, and note targets are mutable public occurrences of those entities.

Separating entities from occurrences avoids indexing identical blob content once per branch. Every candidate must still join to at least one occurrence whose exact ref tip matches the request's current visibility snapshot. Commit-message candidates need a current reachable-ref witness. Note candidates need both a current notes-state witness and a current target witness.

GitHub Blackbird and Zoekt similarly separate content from repository, path, and branch location data. Zoekt represents branch membership independently and requires the trusted browser to constrain search to visible branches.

Sources: GitHub Blackbird architecture, Zoekt design.

Trigrams provide candidates, not matches

Zoekt and Google Code Search index three-character sequences, derive mandatory literals from regular expressions, intersect candidate postings, then verify the complete expression against source content. This avoids scanning every file for ordinary literals and selective regular expressions. Patterns without a mandatory trigram still require bounded scanning.

SQLite FTS5 provides a trigram tokenizer for general substring candidates. Queries shorter than three Unicode characters cannot use that index effectively. A contentless FTS5 table can retain postings without storing another source-text copy. detail=none and columnsize=0 reduce index data, with query-shape restrictions acceptable for generated trigram terms.

FTS5 MATCH, LIKE, and GLOB are not Luigit's public query language. Passing user syntax directly would expose engine syntax and permit scan-prone plans. Luigit must generate bounded internal candidate queries and independently verify each candidate.

Sources: SQLite FTS5 trigram tokenizer and contentless tables, Zoekt design, Google Code Search trigram design.

Rust regular expressions provide bounded verification primitives

Rust regex avoids backreferences and look-around and documents worst-case search proportional to pattern size times searched bytes. Its builder limits compiled size, lazy-DFA cache size, and syntax nesting. The caller must additionally bound pattern bytes, candidate count, per-object bytes, total verified bytes, output, and elapsed work.

regex-syntax can extract required prefix or suffix literal sets from a parsed expression. Extraction may become infinite or inexact for large classes and repetitions. Such output cannot safely narrow candidates unless every possible match retains a proven indexed literal.

The index normalization and exact verifier must implement the same case-insensitive contract. A conservative planner may abandon trigram acceleration and perform a bounded scan whenever Unicode folding cannot produce a complete candidate set. This preserves correctness at the cost of explicit incompleteness under budget.

Sources: Rust regex, RegexBuilder, regex-syntax literal extraction.

SQLite fits the combined search state

Bundled SQLite with FTS5 can live inside the Rust executable through rusqlite. Relational tables can hold entities, occurrences, exact ref tips, coverage, generations, and reclamation state beside contentless trigram postings. One transaction can publish an indexed scope and its coverage marker atomically.

WAL mode allows readers while one indexing writer commits updates. SQLite progress callbacks can interrupt long statement preparation and execution. Application limits can lower SQLite's broad defaults for hostile public input.

Tantivy offers immutable segments, committed reader reloads, delete tombstones, merges, and garbage collection. Its regex query matches indexed terms rather than full original documents. Luigit would still need a separate transactional occurrence and visibility catalog, making Tantivy the larger initial architecture.

Sources: SQLite isolation, SQLite WAL, SQLite query progress callbacks, SQLite limits, rusqlite bundled builds, Tantivy RegexQuery, Tantivy architecture.

Freshness needs semantic reconciliation

A changed Git repository is fully described for indexing purposes by changed public metadata, advertised refs, peeled targets, HEAD, notes tips, or index-policy version. Unchanged object IDs need no content reindexing.

Filesystem events can reduce detection latency but are lossy hints. Inotify queues can overflow, rename correlation is racy, and watched directory structure can change. Git hooks cover only selected mutation paths and would add executable behavior to authoritative repositories. Neither is a complete source of truth.

A short periodic semantic ref scan is the authoritative update detector. A search request also reads current refs before using indexed occurrences. A mismatch suppresses the changed scope immediately, marks it incomplete, and raises its indexing priority. This protects revocation even before background indexing catches up.

Sources: inotify, Git hooks, gix::Repository.

Progressive indexing belongs at ref scope

Each indexed branch, tag, notes namespace, commit-history set, and metadata class needs explicit coverage tied to exact source tips. Coverage states are complete, indexing, skipped, or failed.

On a changed ref, Luigit first stops authorizing its old occurrences. It then compares old and new trees, indexes previously unseen entities, replaces affected occurrences, and publishes the new tip plus coverage in one transaction. Unchanged refs remain searchable.

First build prioritizes repository metadata, refs, the default branch, current notes, other branch and tag snapshots, commit history, then historical note search. Search remains usable for completed scopes and reports missing scopes rather than delaying ordinary navigation.

Sourcegraph's documented repository polling and indexing latency is measured in tens of seconds to minutes, so its operational cadence cannot establish Luigit's seconds-scale target. Its limits and skipped-file reporting remain useful precedents.

Sources: Sourcegraph repository update frequency, Sourcegraph search configuration, Zoekt search limits.

Cache bounds need owned-byte accounting

A cache capacity measured in entry count does not bound memory when values vary from small metadata to large rendered diffs. Admission must use owned bytes plus known overhead, reject individually oversized entries, and separate object, rendered-page, and diff classes. Concurrent identical misses should share one computation.

Positive and negative entries must include the visibility snapshot or be reauthorized on every hit. Timeout, overload, corruption, and uncertain-authorization failures should not become durable negative entries.

Application cache limits are not hard process limits. Concurrent admissions, leased values, allocator overhead, memory maps, and filesystem page cache can exceed an application estimate. A cgroup memory ceiling and a dedicated derived-storage quota provide the hard boundaries.

Sources: Linux cgroup v2, Moka cache semantics, Linux filesystem statistics.

Logical deletion and physical reclamation differ

Removing an occurrence transactionally makes it non-searchable. FTS postings, free SQLite pages, WAL pages, old replacement databases, and open-unlinked files may continue consuming disk. Authorization revocation must not wait for physical reclamation.

FTS5 supports incremental segment merging. SQLite incremental auto-vacuum can return free pages gradually but may increase fragmentation. A full VACUUM can require up to twice the original database size. VACUUM INTO creates a compact consistent snapshot beside the active file but its incomplete output is not crash-safe until validated and published.

Routine maintenance therefore needs bounded merge, checkpoint, and incremental-vacuum work while headroom exists. A side-by-side compacted replacement is exceptional work admitted only when active, staged, temporary, WAL, and emergency-reserve bytes all fit.

Sources: SQLite FTS5 merge controls, SQLite auto-vacuum, SQLite VACUUM, Linux unlink.

Low disk must degrade search before browsing

The derived volume needs high and critical watermarks below its hard quota. Admission must account for current physical allocation, inodes, worst-case new allocation, and an untouchable emergency reserve. Index compaction can require more space before releasing old space, so forcing it after exhaustion can worsen failure.

At the high watermark, Luigit stops optional cache admission and expensive maintenance, evicts caches, removes abandoned temporary files, and performs only already-admissible reclamation. At the critical watermark or after ENOSPC or EDQUOT, it stops index writes and compaction while direct Git browsing continues. Search serves only safely authorized completed coverage and reports degraded or incomplete state.

Crash recovery trusts SQLite transactions, validates schema and integrity before opening search, discards only recognized abandoned derived files, and rebuilds after corruption. Unknown files and paths outside the canonical derived root are never swept.

Sources: SQLite WAL, SQLite VACUUM, Linux rename, Linux fsync, Linux quotas.

Git garbage collection is outside Luigit

Git garbage collection can repack and prune authoritative objects. Notes do not keep annotated objects alive. Concurrent source maintenance may also invalidate an indexing read and require a refreshed repository handle.

Luigit has read-only repository authority. It must never run Git GC, prune, repack, expire reflogs, mutate refs, or otherwise decide source retention. Missing or concurrently replaced objects produce incomplete indexing or a bounded retry. Luigit garbage collection owns only its SQLite search state, memory caches, and recognized temporary derived files.

Sources: Git garbage collection, gix-odb refresh modes.

Conclusions

Use bundled SQLite FTS5 as the v1 search-index candidate. It provides trigram postings, relational provenance, coverage state, atomic updates, one-writer coordination, query interruption, and one-executable packaging through one dependency boundary. Benchmark Tantivy only as the fallback if SQLite misses measured build, size, update, or query targets.

Index valid UTF-8 searchable entities once by stable Git identity. Store no source body solely to produce snippets. Keep exact ref-tip, path, reachability, note-state, and target occurrences separately. Fetch the authoritative object and verify the complete literal or regular expression only after current authorization succeeds.

Generate all FTS candidate queries internally. Use selective mandatory trigrams for literals and regular expressions when extraction is complete. Use a bounded authorized scan for one-character or two-character literals and expressions without safe trigrams. Budget exhaustion returns partial results with a reason; it never silently claims completeness.

Define current authorization as the final successfully observed public metadata and advertised-ref snapshot for the request. Read that snapshot before candidate use and again before output. A changed digest retries once or fails closed. This is the strongest read-only contract available without coordinating repository writers.

Use short periodic semantic ref scans as the freshness authority and filesystem events only to schedule earlier scans. Do not install or execute Git hooks. Changed scopes become unauthorized immediately, then return progressively as exact-tip indexing completes. Unchanged scopes and direct browsing remain available.

Use SQLite's active database for normal incremental commits and bounded reclamation. Use at most one staged compacted replacement, publish it only after validation and durable atomic replacement, and retain the active database until all readers release it. Do not begin a rebuild or compaction without measured temporary-space headroom.

Start without a general persistent cache. Use bounded SQLite and gix caches plus small byte-weighted in-memory rendered-result caches only where measurements justify them. Persistent cache storage can be added later without changing correctness because every value is disposable and reauthorized.

Set independent limits for indexed object bytes, path occurrences, refs, candidates, verified bytes, regular-expression pattern and program size, SQLite work, results, response bytes, search concurrency, indexing concurrency, memory, disk, inodes, and maintenance duration. Host cgroups and filesystem quotas remain the hard limits.

Expose indexed and current snapshot digests, observation time, per-scope coverage, queue age, push-to-search latency, skipped content classes, candidate and verification work, partial-result reasons, database and WAL bytes, free pages, cache weight, evictions, reclamation bytes, failures, and low-disk state. Avoid high-cardinality metric labels for queries, paths, refs, object IDs, or generations.

Rejected conclusions

  • Index contents authorize public output.
  • Reflogs or every local ref belong to the public search corpus.
  • Raw user text may enter FTS MATCH, LIKE, or GLOB syntax.
  • Tantivy regex queries verify expressions against source documents.
  • Short literals or unselective regular expressions justify unbounded scans.
  • Polling or inotify proves an instantaneous concurrent repository state.
  • Git hooks are required for freshness.
  • Cache entry count is a memory bound.
  • Deleting database rows immediately reclaims disk.
  • Full compaction is safe after the volume is already full.
  • Luigit may reclaim authoritative Git objects.

Unresolved questions

  • What exact Luigit policy produces the advertised-ref set, including other refs and hidden-ref configuration?
  • Is request-time final snapshot observation an acceptable definition of current, or must Soft Serve provide coordinated writer publication?
  • What complete Unicode case-insensitive literal contract and normalization will indexing and verification share?
  • Are invalid-UTF-8 and binary blobs excluded from content search or assigned a separate byte-search mode later?
  • What repository, ref, branch, tag, path, blob, commit, and notes-history sizes occur in the deployed corpus?
  • Which maximum object, occurrence, candidate, verification, result, response, memory, and elapsed-work budgets satisfy the p95 targets?
  • Does bundled SQLite FTS5 meet build time, incremental push latency, warm p95, index-size, WAL-size, and peak-memory targets on the recorded corpus?
  • How much temporary disk does FTS merging, incremental vacuum, and side-by-side replacement require on that corpus?
  • Which filesystem and rootless-container quota mechanism enforces both byte and inode limits, and does statvfs report that quota accurately?
  • At what indexed-snapshot lag does Luigit suppress an affected repository rather than return partial search results?
  • Which scopes remain indexed when the configured finite budget cannot hold all required branch, tag, commit-message, current-note, and historical-note coverage?
---
id: BB-RESEARCH-F1ZMINXT
type: research
title: Luigit search indexing freshness caching and garbage collection
---

# Luigit search indexing freshness caching and garbage collection

## Scope

Determine Luigit's searchable representation, update model, cache boundaries, and derived-state reclamation.
Preserve current public reachability, seconds-scale update visibility, bounded resources, and browsing during search failure or rebuild.

## Facts

### Search acceleration is not authorization

Git objects are immutable and content-addressed.
Their public status is mutable because branches, tags, notes refs, hosting visibility, and advertisement policy change.
Force-pushed and deleted-ref objects may remain in packs, reflogs, and old indexes after losing public reachability.

Index membership, object existence, hidden refs, and reflogs cannot authorize output.
This applies to snippets, result counts, facets, suggestions, pagination, and identifiers, not only result links.

Git ref updates are individually atomic.
A concurrent reader may observe only part of a multi-ref transaction.
A read-only consumer cannot know an instantaneous repository state that changes concurrently; it can authorize against a named successfully observed ref snapshot.

Sources: [Git data model](https://raw.githubusercontent.com/git/git/v2.53.0/Documentation/gitdatamodel.adoc), [Git repository layout](https://raw.githubusercontent.com/git/git/v2.53.0/Documentation/gitrepository-layout.adoc), [Git revisions](https://git-scm.com/docs/gitrevisions), [Git update-ref transactions](https://raw.githubusercontent.com/git/git/v2.53.0/Documentation/git-update-ref.adoc).

### Advertised refs need an explicit snapshot

Git protocol v2 `ls-refs` can report refs, symbolic refs, and peeled tags while applying hidden-ref policy.
The repository's complete ref namespace is not necessarily its public advertisement.

A Luigit visibility snapshot must contain the public repository metadata generation, sorted advertised ref names and targets, peeled tag targets, visible `HEAD` symbolic target, selected notes refs, object format, and policy version.
Its digest identifies the authorization observation used by a request and the indexed coverage built from it.

Code reachability and notes-state reachability are separate.
A visible notes ref authorizes its own note history, not the annotated target.
A note result remains visible only when its note state is permitted and its target is reachable from current public code roots.

Sources: [Git protocol v2](https://raw.githubusercontent.com/git/git/v2.53.0/Documentation/gitprotocol-v2.adoc), [Git hidden-ref configuration](https://git-scm.com/docs/git-config#Documentation/git-config.txt-transferhideRefs), [Git notes](https://git-scm.com/docs/git-notes), [Git garbage collection](https://git-scm.com/docs/git-gc).

### Indexed entities and public occurrences are different data

Blob text, commit messages, and note blobs are immutable searchable entities identified by repository, object format, object ID, and kind.
Repository names, refs, paths, branch or tag snapshots, and note targets are mutable public occurrences of those entities.

Separating entities from occurrences avoids indexing identical blob content once per branch.
Every candidate must still join to at least one occurrence whose exact ref tip matches the request's current visibility snapshot.
Commit-message candidates need a current reachable-ref witness.
Note candidates need both a current notes-state witness and a current target witness.

GitHub Blackbird and Zoekt similarly separate content from repository, path, and branch location data.
Zoekt represents branch membership independently and requires the trusted browser to constrain search to visible branches.

Sources: [GitHub Blackbird architecture](https://github.blog/engineering/architecture-optimization/the-technology-behind-githubs-new-code-search/), [Zoekt design](https://github.com/sourcegraph/zoekt/blob/a09cc3cab519ea71490d412b53490f451824e12f/doc/design.md).

### Trigrams provide candidates, not matches

Zoekt and Google Code Search index three-character sequences, derive mandatory literals from regular expressions, intersect candidate postings, then verify the complete expression against source content.
This avoids scanning every file for ordinary literals and selective regular expressions.
Patterns without a mandatory trigram still require bounded scanning.

SQLite FTS5 provides a trigram tokenizer for general substring candidates.
Queries shorter than three Unicode characters cannot use that index effectively.
A contentless FTS5 table can retain postings without storing another source-text copy.
`detail=none` and `columnsize=0` reduce index data, with query-shape restrictions acceptable for generated trigram terms.

FTS5 `MATCH`, `LIKE`, and `GLOB` are not Luigit's public query language.
Passing user syntax directly would expose engine syntax and permit scan-prone plans.
Luigit must generate bounded internal candidate queries and independently verify each candidate.

Sources: [SQLite FTS5 trigram tokenizer and contentless tables](https://www.sqlite.org/fts5.html), [Zoekt design](https://github.com/sourcegraph/zoekt/blob/a09cc3cab519ea71490d412b53490f451824e12f/doc/design.md), [Google Code Search trigram design](https://swtch.com/~rsc/regexp/regexp4.html).

### Rust regular expressions provide bounded verification primitives

Rust `regex` avoids backreferences and look-around and documents worst-case search proportional to pattern size times searched bytes.
Its builder limits compiled size, lazy-DFA cache size, and syntax nesting.
The caller must additionally bound pattern bytes, candidate count, per-object bytes, total verified bytes, output, and elapsed work.

`regex-syntax` can extract required prefix or suffix literal sets from a parsed expression.
Extraction may become infinite or inexact for large classes and repetitions.
Such output cannot safely narrow candidates unless every possible match retains a proven indexed literal.

The index normalization and exact verifier must implement the same case-insensitive contract.
A conservative planner may abandon trigram acceleration and perform a bounded scan whenever Unicode folding cannot produce a complete candidate set.
This preserves correctness at the cost of explicit incompleteness under budget.

Sources: [Rust regex](https://docs.rs/regex/latest/regex/), [RegexBuilder](https://docs.rs/regex/latest/regex/struct.RegexBuilder.html), [regex-syntax literal extraction](https://docs.rs/regex-syntax/latest/regex_syntax/hir/literal/struct.Extractor.html).

### SQLite fits the combined search state

Bundled SQLite with FTS5 can live inside the Rust executable through `rusqlite`.
Relational tables can hold entities, occurrences, exact ref tips, coverage, generations, and reclamation state beside contentless trigram postings.
One transaction can publish an indexed scope and its coverage marker atomically.

WAL mode allows readers while one indexing writer commits updates.
SQLite progress callbacks can interrupt long statement preparation and execution.
Application limits can lower SQLite's broad defaults for hostile public input.

Tantivy offers immutable segments, committed reader reloads, delete tombstones, merges, and garbage collection.
Its regex query matches indexed terms rather than full original documents.
Luigit would still need a separate transactional occurrence and visibility catalog, making Tantivy the larger initial architecture.

Sources: [SQLite isolation](https://sqlite.org/isolation.html), [SQLite WAL](https://www.sqlite.org/wal.html), [SQLite query progress callbacks](https://www.sqlite.org/c3ref/progress_handler.html), [SQLite limits](https://www.sqlite.org/limits.html), [`rusqlite` bundled builds](https://docs.rs/crate/rusqlite/latest/source/README.md), [Tantivy RegexQuery](https://docs.rs/tantivy/latest/tantivy/query/struct.RegexQuery.html), [Tantivy architecture](https://github.com/quickwit-oss/tantivy/blob/main/ARCHITECTURE.md).

### Freshness needs semantic reconciliation

A changed Git repository is fully described for indexing purposes by changed public metadata, advertised refs, peeled targets, `HEAD`, notes tips, or index-policy version.
Unchanged object IDs need no content reindexing.

Filesystem events can reduce detection latency but are lossy hints.
Inotify queues can overflow, rename correlation is racy, and watched directory structure can change.
Git hooks cover only selected mutation paths and would add executable behavior to authoritative repositories.
Neither is a complete source of truth.

A short periodic semantic ref scan is the authoritative update detector.
A search request also reads current refs before using indexed occurrences.
A mismatch suppresses the changed scope immediately, marks it incomplete, and raises its indexing priority.
This protects revocation even before background indexing catches up.

Sources: [inotify](https://man7.org/linux/man-pages/man7/inotify.7.html), [Git hooks](https://raw.githubusercontent.com/git/git/v2.53.0/Documentation/githooks.adoc), [`gix::Repository`](https://docs.rs/gix/latest/gix/struct.Repository.html).

### Progressive indexing belongs at ref scope

Each indexed branch, tag, notes namespace, commit-history set, and metadata class needs explicit coverage tied to exact source tips.
Coverage states are complete, indexing, skipped, or failed.

On a changed ref, Luigit first stops authorizing its old occurrences.
It then compares old and new trees, indexes previously unseen entities, replaces affected occurrences, and publishes the new tip plus coverage in one transaction.
Unchanged refs remain searchable.

First build prioritizes repository metadata, refs, the default branch, current notes, other branch and tag snapshots, commit history, then historical note search.
Search remains usable for completed scopes and reports missing scopes rather than delaying ordinary navigation.

Sourcegraph's documented repository polling and indexing latency is measured in tens of seconds to minutes, so its operational cadence cannot establish Luigit's seconds-scale target.
Its limits and skipped-file reporting remain useful precedents.

Sources: [Sourcegraph repository update frequency](https://sourcegraph.com/docs/admin/repo/update-frequency), [Sourcegraph search configuration](https://sourcegraph.com/docs/admin/search), [Zoekt search limits](https://github.com/sourcegraph/zoekt/blob/a09cc3cab519ea71490d412b53490f451824e12f/api.go).

### Cache bounds need owned-byte accounting

A cache capacity measured in entry count does not bound memory when values vary from small metadata to large rendered diffs.
Admission must use owned bytes plus known overhead, reject individually oversized entries, and separate object, rendered-page, and diff classes.
Concurrent identical misses should share one computation.

Positive and negative entries must include the visibility snapshot or be reauthorized on every hit.
Timeout, overload, corruption, and uncertain-authorization failures should not become durable negative entries.

Application cache limits are not hard process limits.
Concurrent admissions, leased values, allocator overhead, memory maps, and filesystem page cache can exceed an application estimate.
A cgroup memory ceiling and a dedicated derived-storage quota provide the hard boundaries.

Sources: [Linux cgroup v2](https://docs.kernel.org/admin-guide/cgroup-v2.html), [Moka cache semantics](https://docs.rs/moka/latest/moka/), [Linux filesystem statistics](https://man7.org/linux/man-pages/man3/statvfs.3.html).

### Logical deletion and physical reclamation differ

Removing an occurrence transactionally makes it non-searchable.
FTS postings, free SQLite pages, WAL pages, old replacement databases, and open-unlinked files may continue consuming disk.
Authorization revocation must not wait for physical reclamation.

FTS5 supports incremental segment merging.
SQLite incremental auto-vacuum can return free pages gradually but may increase fragmentation.
A full `VACUUM` can require up to twice the original database size.
`VACUUM INTO` creates a compact consistent snapshot beside the active file but its incomplete output is not crash-safe until validated and published.

Routine maintenance therefore needs bounded merge, checkpoint, and incremental-vacuum work while headroom exists.
A side-by-side compacted replacement is exceptional work admitted only when active, staged, temporary, WAL, and emergency-reserve bytes all fit.

Sources: [SQLite FTS5 merge controls](https://www.sqlite.org/fts5.html), [SQLite auto-vacuum](https://www.sqlite.org/pragma.html#pragma_auto_vacuum), [SQLite VACUUM](https://www.sqlite.org/lang_vacuum.html), [Linux unlink](https://man7.org/linux/man-pages/man2/unlink.2).

### Low disk must degrade search before browsing

The derived volume needs high and critical watermarks below its hard quota.
Admission must account for current physical allocation, inodes, worst-case new allocation, and an untouchable emergency reserve.
Index compaction can require more space before releasing old space, so forcing it after exhaustion can worsen failure.

At the high watermark, Luigit stops optional cache admission and expensive maintenance, evicts caches, removes abandoned temporary files, and performs only already-admissible reclamation.
At the critical watermark or after `ENOSPC` or `EDQUOT`, it stops index writes and compaction while direct Git browsing continues.
Search serves only safely authorized completed coverage and reports degraded or incomplete state.

Crash recovery trusts SQLite transactions, validates schema and integrity before opening search, discards only recognized abandoned derived files, and rebuilds after corruption.
Unknown files and paths outside the canonical derived root are never swept.

Sources: [SQLite WAL](https://www.sqlite.org/wal.html), [SQLite VACUUM](https://www.sqlite.org/lang_vacuum.html), [Linux rename](https://man7.org/linux/man-pages/man2/rename.2), [Linux fsync](https://man7.org/linux/man-pages/man2/fsync.2), [Linux quotas](https://docs.kernel.org/filesystems/quota.html).

### Git garbage collection is outside Luigit

Git garbage collection can repack and prune authoritative objects.
Notes do not keep annotated objects alive.
Concurrent source maintenance may also invalidate an indexing read and require a refreshed repository handle.

Luigit has read-only repository authority.
It must never run Git GC, prune, repack, expire reflogs, mutate refs, or otherwise decide source retention.
Missing or concurrently replaced objects produce incomplete indexing or a bounded retry.
Luigit garbage collection owns only its SQLite search state, memory caches, and recognized temporary derived files.

Sources: [Git garbage collection](https://git-scm.com/docs/git-gc), [`gix-odb` refresh modes](https://docs.rs/gix-odb/latest/gix_odb/store/enum.RefreshMode.html).

## Conclusions

Use bundled SQLite FTS5 as the v1 search-index candidate.
It provides trigram postings, relational provenance, coverage state, atomic updates, one-writer coordination, query interruption, and one-executable packaging through one dependency boundary.
Benchmark Tantivy only as the fallback if SQLite misses measured build, size, update, or query targets.

Index valid UTF-8 searchable entities once by stable Git identity.
Store no source body solely to produce snippets.
Keep exact ref-tip, path, reachability, note-state, and target occurrences separately.
Fetch the authoritative object and verify the complete literal or regular expression only after current authorization succeeds.

Generate all FTS candidate queries internally.
Use selective mandatory trigrams for literals and regular expressions when extraction is complete.
Use a bounded authorized scan for one-character or two-character literals and expressions without safe trigrams.
Budget exhaustion returns partial results with a reason; it never silently claims completeness.

Define current authorization as the final successfully observed public metadata and advertised-ref snapshot for the request.
Read that snapshot before candidate use and again before output.
A changed digest retries once or fails closed.
This is the strongest read-only contract available without coordinating repository writers.

Use short periodic semantic ref scans as the freshness authority and filesystem events only to schedule earlier scans.
Do not install or execute Git hooks.
Changed scopes become unauthorized immediately, then return progressively as exact-tip indexing completes.
Unchanged scopes and direct browsing remain available.

Use SQLite's active database for normal incremental commits and bounded reclamation.
Use at most one staged compacted replacement, publish it only after validation and durable atomic replacement, and retain the active database until all readers release it.
Do not begin a rebuild or compaction without measured temporary-space headroom.

Start without a general persistent cache.
Use bounded SQLite and `gix` caches plus small byte-weighted in-memory rendered-result caches only where measurements justify them.
Persistent cache storage can be added later without changing correctness because every value is disposable and reauthorized.

Set independent limits for indexed object bytes, path occurrences, refs, candidates, verified bytes, regular-expression pattern and program size, SQLite work, results, response bytes, search concurrency, indexing concurrency, memory, disk, inodes, and maintenance duration.
Host cgroups and filesystem quotas remain the hard limits.

Expose indexed and current snapshot digests, observation time, per-scope coverage, queue age, push-to-search latency, skipped content classes, candidate and verification work, partial-result reasons, database and WAL bytes, free pages, cache weight, evictions, reclamation bytes, failures, and low-disk state.
Avoid high-cardinality metric labels for queries, paths, refs, object IDs, or generations.

## Rejected conclusions

- Index contents authorize public output.
- Reflogs or every local ref belong to the public search corpus.
- Raw user text may enter FTS `MATCH`, `LIKE`, or `GLOB` syntax.
- Tantivy regex queries verify expressions against source documents.
- Short literals or unselective regular expressions justify unbounded scans.
- Polling or inotify proves an instantaneous concurrent repository state.
- Git hooks are required for freshness.
- Cache entry count is a memory bound.
- Deleting database rows immediately reclaims disk.
- Full compaction is safe after the volume is already full.
- Luigit may reclaim authoritative Git objects.

## Unresolved questions

- What exact Luigit policy produces the advertised-ref set, including other refs and hidden-ref configuration?
- Is request-time final snapshot observation an acceptable definition of current, or must Soft Serve provide coordinated writer publication?
- What complete Unicode case-insensitive literal contract and normalization will indexing and verification share?
- Are invalid-UTF-8 and binary blobs excluded from content search or assigned a separate byte-search mode later?
- What repository, ref, branch, tag, path, blob, commit, and notes-history sizes occur in the deployed corpus?
- Which maximum object, occurrence, candidate, verification, result, response, memory, and elapsed-work budgets satisfy the p95 targets?
- Does bundled SQLite FTS5 meet build time, incremental push latency, warm p95, index-size, WAL-size, and peak-memory targets on the recorded corpus?
- How much temporary disk does FTS merging, incremental vacuum, and side-by-side replacement require on that corpus?
- Which filesystem and rootless-container quota mechanism enforces both byte and inode limits, and does `statvfs` report that quota accurately?
- At what indexed-snapshot lag does Luigit suppress an affected repository rather than return partial search results?
- Which scopes remain indexed when the configured finite budget cannot hold all required branch, tag, commit-message, current-note, and historical-note coverage?