Skip to content

Increase Monitoring Stack Scalability/Resources and Update Ansible - #43

Merged
mamoutou-diarra merged 6 commits into
mainfrom
mamoutou/increase_victoria_resources
Jul 22, 2026
Merged

mamoutou-diarra merged 6 commits into
mainfrom
mamoutou/increase_victoria_resources

Conversation

@mamoutou-diarra

@mamoutou-diarra mamoutou-diarra commented Jul 16, 2026 •

Copy link
Copy Markdown
Collaborator

Context

This PR increases overall scalability of the observability stack to avoid delayed metrics and delayed log insertion. vlstorage (log storage) was running as a single replica with no redundancy and was allocated only 1Gi memory despite real usage of ~7.7Gi. Because there was no redundancy, any brief stall on that single pod halted all log ingestion, surfacing as "all storage nodes unavailable" in vlinsert.
Regarding metrics, a recent experiment used ~51% CPU on the node hosting vmstorage-2 for ~35 minutes, causing i/o timeouts but this has been solved by the other 2 replicas.

Changes

  • Increased vlstorage replicas from 1 to 3 and set real resource requests/limits (it had no limit at all, and its request was 1/7th of its actual usage, which means kubernetes can maintain it on a hoghly overloaded node).
  • Increased vmstorage's memory request (2Gi → 6Gi) to match real usage from kubectl top (~5.5Gi).
  • Set real resource requests/limits for vlselect and vector, which previously had none.
  • Added a new monitoring-critical PriorityClass, applied to all vmetrics/vlogs components, so they're scheduled ahead of and evicted after experiment/simulation pods under resource pressure.
  • Added toleration to vmetrics/vlogs components so that metal-01 becomes a usable scheduling target for them.
  • Added anti-affinity to vlstorage and vmstorage to avoid their replicas to land on the same node.
  • Increased authentik shm to avoid failed to connect to authentik backend: EOF

Other unrelated changes (chore)

  • Created mongodb and dst_dashboard JWT secrets in ansible vault and created corresponding task
  • Updated Ansible cilium task to increase api rate limiter threshold
  • Added portmap in Ansible cilium task to be consistent with running values in the cluster (we enabled portmap directly in the cluster and forgot to update ansible)
  • Moved otlp-collectors to metal-01 node

@mamoutou-diarra mamoutou-diarra self-assigned this Jul 16, 2026
@mamoutou-diarra mamoutou-diarra added the ift IFT commitments label Jul 16, 2026
@mamoutou-diarra mamoutou-diarra changed the title Increase Monitoring Stack Scalability and Resources Increase Monitoring Stack Scalability/Resources and Update Ansible Jul 17, 2026
@mamoutou-diarra mamoutou-diarra linked an issue Jul 17, 2026 that may be closed by this pull request
@mamoutou-diarra
mamoutou-diarra merged commit 692f110 into main Jul 22, 2026
@github-project-automation github-project-automation Bot moved this to Done in DST Jul 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ift IFT commitments

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Analyze current stack (recurring)

2 participants