Files
twenty/packages
Weiko 024b379e1a Instrument Node runtime and workspace cache metrics (#23107)
## Context

High API tail latency can come from either downstream work or the
Node.js process itself being unable to schedule work. Existing traces
expose database and HTTP spans, but they do not provide a continuous
event-loop signal or identify time spent rebuilding individual workspace
metadata cache entries.

## What changed

- Enable Sentry's built-in Node runtime integration with a 30-second
collection interval.
- Collect only event-loop delay p99, event-loop delay max, and
event-loop utilization. CPU, memory, p50, and uptime metrics remain
disabled because existing infrastructure telemetry already covers those
areas.
- Add a parent span around workspace metadata cache invalidation and
recomputation.
- Add child spans around cache-provider computation, including the cache
key, recomputation strategy, and whether the provider uses local data
only.

## Telemetry scope

- Cache hits do not create spans.
- Cache spans use `onlyIfParent`, so they are recorded only inside an
already-sampled trace.
- Runtime metrics are three low-cardinality values every 30 seconds per
server process.
- This does not add Prometheus histograms, per-cache-key metric labels,
database pool gauges, or a custom runtime collector.
- No cache behavior or invalidation semantics change.

This should let us distinguish event-loop stalls from downstream
latency, then identify which cache provider contributes to a slow cache
rebuild without materially increasing telemetry volume.
2026-07-21 14:02:58 +00:00
..
2026-07-21 15:48:30 +02:00