5f94ee3e02
## Context The per-pod in-process workspace metadata cache (`WorkspaceCacheService`) evicts by a fixed **1,000-entry count**. Each workspace's cached metadata is ~1 MB (dominated by the flat `field-metadata` map) over ~10–13 entries, so 1,000 entries ≈ only a few dozen workspaces per pod. On a multi-tenant instance with far more active workspaces, the L1 cache thrashes — LRU-evicting and re-fetching the ~1 MB of maps from Redis on misses — and since the cache sits in one AZ while pods span both, ~half of that transfer is billed cross-AZ. (In prod this cache node serves ~2.6 TB/day.) ## What this does - **Raise `MAX_LOCAL_CACHE_ENTRIES` 1,000 → 7,500** (~500 workspaces at ~1 MB each; server pods are 4 GiB / `--max-old-space-size=3500`, so this stays well within the heap). - **Add a `workspace-metadata-cache/local-eviction` counter** (incremented by the number of entries dropped each time the cache hits capacity) so we can see capacity-driven evictions in metrics and tune the limit from real data rather than guessing. Eviction stays **batched** (`MIN_EVICT_KEYS`), so the sort runs about once per 100 inserts at steady state rather than on every write. ### Why count, not bytes An earlier iteration bounded by measured bytes, but that required `JSON.stringify`-ing every cached value (incl. the ~1 MB field-metadata maps) on every write — meaningful CPU/GC overhead on the fill path. A raised count cap avoids that entirely; the new eviction metric gives us the signal to right-size it. No change to cache semantics, hashing, or the Redis format.