Commit Graph

17 Commits

Author SHA1 Message Date
Charles Bochet 1ebcbdda42 feat(server): export event-loop delay and workspace-cache recompute metrics (#23751)
## What

Exports three sets of metrics to Prometheus to diagnose the recurring
"Slow DB Query" Sentry issues on `POST /graphql` (e.g. the
`fieldMetadata` select):

- `twenty_nodejs_eventloop_delay_seconds` (mean/p50/p99/max) +
`twenty_nodejs_eventloop_utilization`
- `twenty_workspace_cache_recompute_duration_seconds{cache_key}` — wall
time per provider `computeForCache`
- `twenty_workspace_cache_redis_write_duration_seconds` — serialize +
Redis write time for recomputed entries

## Why

Investigation of these issues showed:
- The flagged query executes in ~1ms (prod EXPLAIN), so it is not a
query/index problem.
- Event counts do **not** correlate with connection-pool acquire latency
(Pearson ~0 against both p99 and the direct count of >1s acquires), so
it is not pool contention.
- Event counts **do** correlate with pod CPU (Pearson +0.40).

The leading explanation is that the slow `db` span is inflated by
event-loop saturation during the workspace metadata cache recompute:
`Promise.all` parallelizes the I/O, but the synchronous work it cannot
parallelize (TypeORM entity hydration of JSONB-heavy result sets, then
`JSON.stringify` of the flat-map payloads into Redis) blocks the single
event-loop thread, so an awaiting query resolves ~1.3s late.

Node event-loop delay was only being collected by Sentry's
`nodeRuntimeMetricsIntegration`, never exported to Prometheus, so it
could not be graphed or correlated in Grafana. These metrics confirm (or
refute) the mechanism and give a before/after baseline for the fix.
Grafana panels land in a companion twenty-infra PR.

## Notes

- Metric-only change; no behavioral change to the cache.
- Uses the same OTel `MetricsService` / meter as the existing DB pool
metrics.


<!-- This is an auto-generated description by cubic. -->
<a
href="https://cubic.dev/pr/twentyhq/twenty/pull/23751?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>
<!-- End of auto-generated description by cubic. -->
2026-08-04 16:50:50 +02:00
Weiko fc6a95a37f Throttle local cache expiration sweeps (#23579)
## Context

The workspace cache and core entity cache keep bounded in-process maps.
Entries that have not been read for 30 minutes are removed by an
expiration sweep.

Before this PR, every cache read synchronously walked the entire local
cache, including every stored version, before performing the actual
lookup. The cost therefore grew with the number of cached entries even
when there was nothing to expire. Both caches are used on common server
request paths, so these repeated full-map scans add unnecessary CPU work
and short-lived allocations, which can contribute to event-loop and
garbage-collection pressure under load.

## What changed

- Run each cache's expiration sweep at most once per minute.
- Keep the existing expiration logic unchanged and use one captured
timestamp for the complete sweep.
- Cover both cache services with tests proving that repeated reads
within the interval trigger one sweep and that sweeping resumes after
the interval.

Normal cache reads now pay only for a timestamp check and branch. The
full `O(cache size)` scan runs at most once per minute per process.

## Safety

This does not change cache freshness:

- The 100 ms local freshness window and Redis hash validation still run
as before.
- Explicit cache invalidation is unchanged.
- The 30-minute inactivity threshold is unchanged.
- LRU eviction still runs when entries are inserted.
- Existing local-cache size limits remain unchanged.

An unused entry can remain in memory for at most one additional minute
before the next sweep. This may marginally increase average retained
memory, but it cannot cause unbounded growth or allow stale data to
bypass the existing hash validation.

## Expected impact

This removes a cache-size-dependent operation from a high-frequency
path. The expected benefit is lower CPU and allocation overhead, less
garbage-collection pressure, and improved tail latency when local caches
are populated.

This is intentionally a narrow optimization. It does not claim to
address every source of API tail latency.

## Validation

- `yarn nx typecheck twenty-server`
- Type-aware Oxlint on the changed files
- Oxfmt on the changed files
- Targeted workspace-cache and core-entity-cache Jest suites, 23 tests
passing
2026-07-30 14:24:16 +00:00
Weiko 0108a34765 Rekey metadata caches when flat map hashes change (#23164)
## Context

The `/metadata` GraphQL response cache (`ObjectMetadataItems`,
`FindAllViews`) and the workspace SDL cache were keyed on
`workspace.metadataVersion`, an integer bumped on every object/field
migration. That mechanism is legacy (the migration runner literally
calls it `getLegacyCacheInvalidationPromises`): the data plane already
moved to `WorkspaceCacheService`, which versions each flat entity map
with its own hash minted on invalidation.

Version keying had two concrete costs: the version bump was the only
proactive invalidation for `ObjectMetadataItems`, and keeping
`FindAllViews` fresh required `flushGraphQLOperation`, a full Redis
keyspace SCAN on every relevant migration. It is also the main blocker
for deprecating `metadataVersion` entirely.

This PR re-keys both caches on the flat-map hashes instead.

## What changed

**Response cache** (`use-cached-metadata.ts`): the key is now
`{operation}:{workspaceId}:{combinedDependencyHash}[:{userWorkspaceId}]:{locale}:{queryHash}`.
Each cached operation declares which flat maps its resolvers read
(`metadata-graphql-operations-to-cache.constant.ts`) plus a scope:
`ObjectMetadataItems` stays workspace-shared, `FindAllViews` is per-user
because unlisted-view visibility depends on the caller. When any
declared map changes, its hash rotates and the key rotates with it; no
flush needed. The key is resolved once per request and reused in
`onResponse`, so a rotation mid-request can never cache a response under
a fresher key than the data it was built from. If hash resolution fails,
the request is served uncached (Sentry-captured).

This also fixes three pre-existing key soundness gaps: `FindAllViews`
ignored locale although view names are translated server-side, the query
hash ignored GraphQL variables (`$viewTypes`), and mid-request version
rotation could re-key between request and response.

**SDL cache** (`workspace-graphql-schema-sdl.service.ts`): keyed on the
combined hash of the four maps the schema is generated from, taken from
the same `getOrRecomputeWithHashes` call that returns the data, so key
and content cannot skew. The `metadataVersion` read/seed block there is
gone; the Redis seed moved to `middleware.service.ts` so the
`X-Schema-Version` "refresh the page" check keeps working after the
Redis key's TTL expires.

**`WorkspaceCacheService`**: the internal pipeline now threads `{data,
hashes}` through every stage (local hit, hash validation, Redis fetch,
recompute) and the memoizer stores the pair, so returned hashes are
always consistent with returned data. New public
`getOrRecomputeWithHashes` and `getOrRecomputeCombinedHash`
(hashes-first: one MGET of the small `:hash` keys, full pipeline only
for missing ones, so cold pods never pull map payloads just to build a
key).

**Atomic pair writes** (`cache-storage.service.ts`): `mset` on the Redis
driver now delegates to the store's own `mset` (a MULTI of `SET ... PX`,
or native `MSET`), grouped by TTL. Previously it was a `Promise.all` of
independent SETs, so two concurrent recomputes could interleave and
leave one recompute's `:data` next to the other's `:hash`; with
hash-keyed caches that torn pair could persist a stale response under a
live key. `CoreEntityCacheService` writes through the same method and is
fixed for free.

**Cleanup**: `flush()` lost its `metadataVersion` parameter (always
pattern-flush per key on workspace deletion),
`METADATA_VERSIONED_WORKSPACE_CACHE_KEY` became
`HASH_KEYED_WORKSPACE_CACHE_KEYS` with the `MetadataVersion` key
relocated to `WORKSPACE_CACHE_KEYS` and the dead `ORMEntitySchemas`
entry removed.

## Deliberately unchanged

- `incrementMetadataVersion` and all its callers stay: the version still
feeds the `X-Schema-Version` check and the pinned upgrade commands.
Deprecating the column is a later stage.
- The runner's `FindAllViews` pattern-flush is kept for exactly one
release: view-only migrations never bump `metadataVersion`, so old pods
in a rolling deploy have no other invalidation signal for their
version-keyed entries. It gets deleted next release, which removes the
SCAN entirely.
- Old-shape cache entries are not migrated; they expire via the 7-day
TTL.

## Known limitations (follow-ups, not regressions)

- The plugin reads dependency hashes Redis-fresh while resolvers can
serve up to 10s-old memoized data, so a request landing right after a
migration can cache a pre-rotation response under the new key. Same
shape existed under `metadataVersion`; closing it needs request-scoped
snapshot plumbing.
- Concurrent recomputes are last-writer-wins (lost update). Fencing with
a conditional write is a follow-up.

## Validation

- Unit: response-cache plugin behavior (scope, key stash, serve-uncached
on failure, prototype-name guard), atomic `mset` batching, existing
`WorkspaceCacheService` spec passing unchanged.
- Integration: a new drift-guard spec runs the real
`ObjectMetadataItems`/`FindAllViews` operations with full frontend
selection sets against the in-process app, spies on
`WorkspaceCacheService`, and fails if resolvers read a flat map missing
from the declared dependency lists, so the constant cannot silently
drift.
- Manual against a live server: creating a field rotates the field-map
hash and the very next `ObjectMetadataItems` response contains it (hash
rotation is now its only invalidation path); warm hits are ~5ms; SDL
entries appear under hash-shaped keys via introspection.

## Suggested reading order

1. `workspace-cache.service.ts`, `workspace-cache-key.type.ts`,
`combine-cache-hashes.util.ts` (the `{data, hashes}` pipeline)
2. `use-cached-metadata.ts`,
`metadata-graphql-operations-to-cache.constant.ts`,
`metadata.module-factory.ts` (response cache)
3. `workspace-graphql-schema-sdl.service.ts`,
`workspace-cache-storage.service.ts` (SDL cache and renames)
4. `middleware.service.ts` (metadata version seed relocation)
5. `cache-storage.service.ts` (atomic writes)
6. Tests
2026-07-22 16:50:09 +00:00
Weiko 024b379e1a Instrument Node runtime and workspace cache metrics (#23107)
## Context

High API tail latency can come from either downstream work or the
Node.js process itself being unable to schedule work. Existing traces
expose database and HTTP spans, but they do not provide a continuous
event-loop signal or identify time spent rebuilding individual workspace
metadata cache entries.

## What changed

- Enable Sentry's built-in Node runtime integration with a 30-second
collection interval.
- Collect only event-loop delay p99, event-loop delay max, and
event-loop utilization. CPU, memory, p50, and uptime metrics remain
disabled because existing infrastructure telemetry already covers those
areas.
- Add a parent span around workspace metadata cache invalidation and
recomputation.
- Add child spans around cache-provider computation, including the cache
key, recomputation strategy, and whether the provider uses local data
only.

## Telemetry scope

- Cache hits do not create spans.
- Cache spans use `onlyIfParent`, so they are recorded only inside an
already-sampled trace.
- Runtime metrics are three low-cardinality values every 30 seconds per
server process.
- This does not add Prometheus histograms, per-cache-key metric labels,
database pool gauges, or a custom runtime collector.
- No cache behavior or invalidation semantics change.

This should let us distinguish event-loop stalls from downstream
latency, then identify which cache provider contributes to a slow cache
rebuild without materially increasing telemetry volume.
2026-07-21 14:02:58 +00:00
Félix Malfait d4ac6e752b fix(server): stop cross-pod recompute cascade on localDataOnly workspace cache keys (#22980)
## Context

Prod investigation (Sentry, last 7 days) traced the current slowness to
the per-pod workspace cache. Every recompute of a `localDataOnly` key
(`ORMEntityMetadatas`, `flatWorkspaceMemberMaps`) published a fresh
`crypto.randomUUID()` as the shared Redis validation hash. Because these
keys recover from a hash mismatch by recomputing (their data never
enters Redis), one miss on one pod invalidated the local copy on every
other pod; each of their recomputes minted yet another hash,
re-invalidating everyone else. The fleet never converges.

Measured impact in prod:
- The `ORMEntityMetadatas` rebuild (full `objectMetadata` +
`fieldMetadata` + `application` queries, ~220ms combined, plus
`EntityMetadataBuilder.build`) ran **~963k times in 24h** (~11/s),
roughly 58h of cumulative Postgres time per day.
- The hottest single workspace recomputed its schema metadata 51k
times/day (once per 1.7s).
- Second-order effects: `POST /metadata` averaged 26.6s (p95 2.3s, so a
tail hangs for minutes on pool/event-loop starvation), GraphQL p95 went
846ms (v2.20.0) to 1744ms (v2.21.0), `Query read timeout` on trivial
cron queries at 18x baseline.

The random hash was correct in the original design (#15962): it is a
generation token, and Redis-backed keys recover absorptively by adopting
hash+data from Redis. #16287 added `localOnly` keys (EntityMetadata[] is
not serializable) whose recovery is generative, which silently broke the
invariant later documented in #18649 ("hashes change only on
invalidateAndRecompute").

## What this does

- Recovery recomputes now **adopt** the hash already present in Redis
instead of minting a new one, and write nothing back. A miss costs one
recompute on one pod instead of an unbounded fleet-wide loop.
- Minting is reserved for `invalidateAndRecompute` (real metadata
changes, propagation semantics unchanged, including the frontend
collectionHashes contract) and the bootstrap case where Redis has no
hash.
- The bootstrap write uses **SET NX** (new
`CacheStorageService.setIfAbsent`) instead of a plain overwrite: a slow
bootstrap recompute could otherwise land after a concurrent
`invalidateAndRecompute` mint and clobber it with a hash of
pre-migration data. Under the old code that clobber self-healed via the
cascade; with adopt semantics it would pin stale data, so the bootstrap
write must lose that race. A losing pod keeps its result locally as
provisional and converges on the winning hash at its next revalidation
(covered by a dedicated race test).

Redis-backed keys are untouched: same fetch-on-mismatch recovery, same
mint-and-write on `missingInRedis`.

## Expected effect and how to verify

`FieldMetadataEntity`/`ObjectMetadataEntity`/`ApplicationEntity`
full-workspace query counts in Sentry should collapse from ~1M/day to
the true metadata-change rate, and with them the DB pool pressure behind
the `/metadata` latency tail. This also makes local-cache eviction
(`MAX_LOCAL_CACHE_ENTRIES`, #22946) cheap: the cap can be tuned purely
for RAM.

Complementary to, not competing with, the planned Redis pub/sub
invalidation: a version token in Redis is still needed for restart
catch-up, and this PR gives it sound semantics.

---
_Generated by [Claude
Code](https://claude.ai/code/session_01T3JUHwXJHPmZDZTrv6YTDi)_

<!-- This is an auto-generated description by cubic. -->
<a
href="https://cubic.dev/pr/twentyhq/twenty/pull/22980?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>
<!-- End of auto-generated description by cubic. -->
2026-07-20 15:26:49 +02:00
Charles Bochet 9860871eac perf(server): lower workspace local cache cap to 6,000 entries (#22946)
## Context

#21954 raised `MAX_LOCAL_CACHE_ENTRIES` from 1,000 to 7,500 to stop the
per-pod workspace metadata cache from thrashing against Redis. That
worked: the cache node's egress dropped from ~1.1-1.8 TB/day to ~0.5-0.7
TB/day and has stayed there.

The memory side of the trade turned out tighter than the sizing assumed.
Prod-eu metrics over the past month:

- Average server pod working set rose from ~1.9-2.4 GiB (before the
rollout on 06/26) to ~3.0-3.4 GiB, with individual pods peaking at
3.5-3.6 GiB, right at the `--max-old-space-size=3500` heap ceiling on 4
GiB pods.
- The `workspace-metadata-cache/local-eviction` counter added in #21954
stayed at zero for three weeks, then on 07/15 between 08:30 and 09:30
UTC all 7 server pods hit the 7,500-entry cap within an hour (no deploy
that morning; organic growth crossed the threshold during the European
morning peak). Since then the fleet evicts continuously (~167k
entries/day).
- Server pods were terminated with reason OOMKilled on 8 days of the
past month (1-2 pods each time).

The per-entry payload also grew since the cap was sized: search field
metadata backfills, two new per-workspace cache components
(`flatSearchFieldMetadataMaps`, `workflowAutomatedTriggerMaps`), and the
heap-only `ORMEntityMetadatas` entries. So the byte ceiling implied by
7,500 entries now sits exactly at the heap limit instead of comfortably
under it.

## What this does

Lower `MAX_LOCAL_CACHE_ENTRIES` 7,500 → 6,000. At the measured ~170-200
KB average entry size this frees roughly 300-500 MB of steady-state heap
per pod, restoring headroom under the 3.5 GB old-space limit.

Traffic cost should be small: eviction churn at the 7,500 cap only moved
cache egress from ~0.55 to ~0.65-0.68 TB/day, so a 20% smaller cap
should keep the bulk of the #21954 reduction. The eviction counter gives
us the feedback loop; if egress climbs materially we can split the
difference, and if pods still run hot the next lever is byte-aware
eviction for the few known-heavy keys.
2026-07-16 12:34:02 +02:00
Charles Bochet 5f94ee3e02 perf(server): raise workspace local cache size and meter evictions (#21954)
## Context

The per-pod in-process workspace metadata cache
(`WorkspaceCacheService`) evicts by a fixed **1,000-entry count**. Each
workspace's cached metadata is ~1 MB (dominated by the flat
`field-metadata` map) over ~10–13 entries, so 1,000 entries ≈ only a few
dozen workspaces per pod. On a multi-tenant instance with far more
active workspaces, the L1 cache thrashes — LRU-evicting and re-fetching
the ~1 MB of maps from Redis on misses — and since the cache sits in one
AZ while pods span both, ~half of that transfer is billed cross-AZ. (In
prod this cache node serves ~2.6 TB/day.)

## What this does

- **Raise `MAX_LOCAL_CACHE_ENTRIES` 1,000 → 7,500** (~500 workspaces at
~1 MB each; server pods are 4 GiB / `--max-old-space-size=3500`, so this
stays well within the heap).
- **Add a `workspace-metadata-cache/local-eviction` counter**
(incremented by the number of entries dropped each time the cache hits
capacity) so we can see capacity-driven evictions in metrics and tune
the limit from real data rather than guessing.

Eviction stays **batched** (`MIN_EVICT_KEYS`), so the sort runs about
once per 100 inserts at steady state rather than on every write.

### Why count, not bytes
An earlier iteration bounded by measured bytes, but that required
`JSON.stringify`-ing every cached value (incl. the ~1 MB field-metadata
maps) on every write — meaningful CPU/GC overhead on the fill path. A
raised count cap avoids that entirely; the new eviction metric gives us
the signal to right-size it.

No change to cache semantics, hashing, or the Redis format.
2026-06-22 17:00:29 +02:00
Etienne 3306d66f5b cleaning - remove logs (#19445) 2026-04-08 13:05:47 +00:00
Etienne 579714b62f Investigate memory leak (#19213)
In search for cache leak + Remove old instrumentation
2026-04-02 11:23:10 +00:00
Charles Bochet 191a277ddf fix: invalidate rolesPermissions cache + add Docker Hub auth to CI (#19044)
## Summary

### Cache invalidation fix
- After migrating object/field permissions to syncable entities (#18609,
#18751, #18567), changes to `flatObjectPermissionMaps`,
`flatFieldPermissionMaps`, or `flatPermissionFlagMaps` no longer
triggered `rolesPermissions` cache invalidation
- This caused stale permission data to be served, leading to flaky
`permissions-on-relations` integration tests and potentially incorrect
permission enforcement in production after object permission upserts
- Adds the three permission-related flat map keys to the condition that
triggers `rolesPermissions` cache recomputation in
`WorkspaceMigrationRunnerService.getLegacyCacheInvalidationPromises`
- Clears memoizer after recomputation to prevent concurrent
`getOrRecompute` calls from caching stale data

### Docker Hub rate limit fix
- CI service containers (postgres, redis, clickhouse) and `docker
run`/`docker build` steps were pulling from Docker Hub
**unauthenticated**, hitting the 100-pull-per-6-hour rate limit on
shared GitHub-hosted runner IPs
- Adds `credentials` blocks to all service container definitions and
`docker/login-action` steps before `docker run`/`docker compose`
commands
- Uses `vars.DOCKERHUB_USERNAME` + `secrets.DOCKERHUB_PASSWORD`
(matching the existing twenty-infra convention)
- Affected workflows: ci-server, ci-merge-queue, ci-breaking-changes,
ci-zapier, ci-sdk, ci-create-app-e2e, ci-website,
ci-test-docker-compose, preview-env-keepalive, spawn-twenty-docker-image
action
2026-03-27 17:32:53 +01:00
Charles Bochet 06efee1eef feat: hash-based metadata staleness detection (#18649)
## Summary

Replace the single `metadataVersion` integer with per-entity-type
**collection hashes** for granular metadata staleness detection. The
backend already generates a UUID per flat entity map on each cache
recompute (`crypto.randomUUID()` in `WorkspaceCacheService`); we now
expose these via the minimal metadata endpoint and SSE events so the
frontend can compare and know exactly which entity types are stale.

### Key changes

**Backend:**
- `WorkspaceCacheService.getCacheHashes()` — new public method that
reads only `:hash` keys from Redis without fetching full data
- `MinimalMetadataDTO` — added `collectionHashes: Record<string,
string>` (JSON scalar mapping `AllMetadataName` → collection hash),
removed `metadataVersion`
- `MetadataEventDTO` — added optional `updatedCollectionHash` field to
SSE events
- `MetadataEventsToDbListener` — reads the collection hash for the
affected entity type after cache invalidation and attaches it to the SSE
event before publishing
- `MinimalMetadataService` — no longer queries the workspace table; uses
`getCacheHashes()` for all flat entity maps and maps cache keys to
`AllMetadataName` locally

**Frontend:**
- `metadataCollectionHashesState` — new Jotai atom with
`atomWithStorage` + `getOnInit: true` storing
`Partial<Record<MetadataEntityKey, string>>`
- `mapAllMetadataNameToEntityKey()` — explicit mapping from backend
`AllMetadataName` to frontend `MetadataEntityKey` (23 entries)
- `useLoadMinimalMetadata` — stores `collectionHashes` from server,
computes `staleEntityKeys` by comparing local vs server hashes
- `patchMetadataStoreFromSSEEvent()` — accepts optional
`updatedCollectionHash` and updates `metadataCollectionHashesState`
- All 11 SSE effect components — pass
`eventDetail.updatedCollectionHash` through to the patch function
- `useStaleMetadataEntities` — new hook returning entity keys missing
from collection hashes (not yet loaded/synced)
- `resetMetadataStore()` — also clears collection hashes
- Deleted `metadataVersionState` (superseded by collection hashes)

### Design decisions

- **No change to hash generation** — existing `crypto.randomUUID()` is
sufficient. Hashes are persisted in Redis, survive server restarts, and
change only on `invalidateAndRecompute`.
- **"Collection hash" naming** — used consistently to clarify the hash
represents an entire entity collection (e.g., all views), not a single
record.
- **Mapping localized** — backend `WorkspaceCacheKeyName` →
`AllMetadataName` mapping lives in the minimal metadata service.
Frontend `AllMetadataName` → `MetadataEntityKey` mapping lives in a
local utility. Nothing in `twenty-shared`.
- **Backward compatible** — `collectionHashes` is additive;
`updatedCollectionHash` is nullable.
2026-03-14 23:38:37 +01:00
Félix Malfait 414f68fb63 fix(workspace-cache): memory leak in deleteFromLocalCache (#17686)
## Problem

The `deleteFromLocalCache` method was only setting `lastHashCheckedAt=0`
instead of actually deleting the cache entry. This caused old versions
to accumulate in the local cache when `invalidateAndRecompute` was
called.

### Flow that causes the leak:

1. `invalidateAndRecompute()` is called
2. `flush()` → `deleteFromLocalCache()` — **only sets
`lastHashCheckedAt=0`, keeps old data**
3. `recomputeDataFromProvider()` → `setInLocalCache()` — **adds new
version with new hash**
4. `cleanupStaleVersions()` — **never called** (only triggered from
`getFromLocalCache` path)

Result: Each `invalidateAndRecompute` call adds a new version to
`entry.versions` without removing the old one.

### Impact

For migration commands (like
`1-17-migrate-attachment-to-morph-relations`) that process thousands of
workspaces, this caused significant memory growth:
- Each workspace calls `getOrRecompute` (version 1)
- Then calls `invalidateAndRecompute` (version 2 added, version 1 stays)
- Memory accumulates as the command processes more workspaces

## Solution

Actually delete the local cache entry in `deleteFromLocalCache`, so
`recomputeDataFromProvider` starts fresh.

```typescript
// Before (memory leak)
if (isDefined(entry)) {
  entry.lastHashCheckedAt = 0;
}

// After (proper cleanup)
this.localCache.delete(localKey);
```

---------

Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
2026-02-04 09:47:45 +00:00
Weiko bd8ed03990 Add TTL eviction to local data cache (#16510)
We keep different versions of our cache to avoid race conditions but we
never evict stale data. This PR should fix that
2025-12-12 10:14:46 +01:00
Weiko 3d95d6ca00 Add local only cache to cache service and cache typeorm entity metadata (#16287)
## Problem

buildEntityMetadatas in GlobalWorkspaceOrmManager is computationally
expensive and was running on every executeInWorkspaceContext call. This
method uses
TypeORM's EntitySchemaTransformer and EntityMetadataBuilder to build
metadata for all workspace entities (30-50+ objects with many fields
each).

The resulting EntityMetadata[] is not serialisable which means it cannot
be cached in Redis because they contain:
- Circular references
- Functions/methods
- References to the DataSource instance

## Solution

Extended the workspace cache system to support local-only caching, then
created a cache provider for entityMetadatas.

## Implementation details
Updated @WorkspaceCache decorator (workspace-cache.decorator.ts)
- Added localOnly?: boolean option to skip Redis storage for
non-serializable data
Created WorkspaceEntityMetadatasCacheService
- Computes entity metadatas from DB to avoid race condition, this is
acceptable
Simplified GlobalWorkspaceOrmManager
- Now fetches entityMetadatas from cache instead of rebuilding on every
call
Updated Workspace migration runner - the only entry point where metadata
can change
- Now invalidate the new 'entityMetadata' local cache when
shouldIncrementMetadataGraphqlSchemaVersion is true (== field/object
mutations)
2025-12-03 19:50:40 +01:00
Weiko 13e283fc3a Rename roleTargets -> roleTarget (#16247) 2025-12-02 14:39:54 +01:00
Weiko 1eb2e44058 Refactor workspace cache service (#16208)
## Context
We've recently introduced a new workspace cache service which now acts
as a cache access and local storage for all workspace related data,
deprecating the individual specific services.
- Better performance through multiple caching/fetching strategies
- Consistent data access patterns across the codebase
- Reduced redis queries through MGET/MSET/PIPELINE with multiple cache
keys
2025-12-01 17:08:21 +01:00
Weiko 0a2d42e79f Implement workspace cache storage (#15962)
## Context
Implementing a single service managing all the cache scoped to a
workspace, this will be dynamically injected as a WorkspaceContext in
the app during a request lifetime (ingested by the future global
datasource for example).

Usage:

```typescript
this.globalWorkspaceOrmManager.executeInWorkspaceContext(
        authContext,
        async () => {
           // Everything here will have access to a workspaceContext, containing all the cache data + a ready to use datasource with workspace scoped metadata, permissions, feature flags, etc...
           // Internally will call loadWorkspaceContext
        }
```
Note: executeInWorkspaceContext will probably be owned by a higher level
service later and not only the ORM.

```typescript
  private async loadWorkspaceContext(
    authContext: WorkspaceAuthContext,
  ): Promise<WorkspaceContext> {
    const workspaceId = authContext.workspace.id;

    const cache =
      await this.workspaceContextCacheService.get<WorkspaceContextData>(
        workspaceId,
        [
          'objectMetadataMaps',
          'metadataVersion',
          'featureFlagsMap',
          'permissionsPerRoleId',
        ],
      );

    return {
      authContext,
      objectMetadataMaps: cache.objectMetadataMaps,
      metadataVersion: cache.metadataVersion,
      featureFlagsMap: cache.featureFlagsMap,
      permissionsPerRoleId: cache.permissionsPerRoleId,
    };
  }
  ```
  
  The cache retrieval strategy is as followed:
```
- Check if there is an ongoing promise fetching data from the cache =>
return the promise.
- Check in the local cache entry if lastCheckedAt has expired. If not,
return as it is without querying redis.
- Check in redis the cache entry hash and compare with local cache entry
hash, if they are the same return the local cache entry data
- Check in redis the cache entry data, if it's there return it and store
it into the local cache entry data and update local cache entry hash. If
it's not there recompute the data by querying the DB and update both
redis and local cache
```
2025-11-21 12:44:24 +01:00