From 400,000 to over 1 million metrics: How to break the scaling barrier and reduce operational costs

How a leading retail organisation broke through vertical scaling limits, achieving resilient monitoring for more than one million metrics while reducing operational costs.

Initial situation and context

A global retail giant ran its critical infrastructure on a Cloud Native architecture based on Google Kubernetes Engine (GKE) and microservices. Its observability ecosystem centralised alert management in a single, large Prometheus instance.

Exponential business growth pushed this architecture to its limits. With more than 400,000 active metric time series in its primary cluster, the organisation faced vertical scaling constraints—both physical and financial. This reliance on a single instance represented a critical risk: any failure of the primary replica would result in a complete loss of operational visibility and potential service disruption.

What was the need?

To ensure the stability of its digital operations and optimise resources, the company needed to evolve towards a distributed architecture. The key objectives were:

  • Future scalability: Prepare the infrastructure to support projected growth beyond 1 million metrics.
  • Resource efficiency: Eliminate Kubernetes node capacity waste caused by oversizing a single instance.
  • Data flexibility: Enable management of different long-term data retention policies and reduce reliance on costly logs in favour of lighter-weight metrics.

How the solution was implemented

Datadope led the transition from a monolithic model to a horizontally scalable architecture across more than 10 Kubernetes clusters. The technical solution included:

  • Prometheus + Thanos architecture: Deployment of multiple Prometheus instances combined with the full Thanos stack (Sidecar, Compactor and Store). This decentralised ingestion while maintaining a unified global view and efficient long-term storage.
  • Acceleration with Redis: Integration of Redis for time-series caching, improving query performance and reducing search latency.
  • Operational standardisation: Definition of new labelling policies and rollout of a self-service alert management system, empowering development teams while ensuring data governance.

Results and benefits

The transformation of the observability architecture delivered an immediate impact on the organisation’s technical and economic efficiency:

  • Direct infrastructure savings: Resource optimisation freed up nearly 10 dedicated Kubernetes nodes, delivering monthly savings of several thousand euros.
  • Improved performance: Data query times fell by 15%, while Prometheus memory consumption decreased by 10%.
  • Increased capacity: The system was validated to scale beyond 1 million metrics, tripling its original capacity without compromising stability.
  • SaaS licensing efficiency: Increased metric capacity enables reduced log ingestion in third-party platforms, with an estimated potential saving of 10%–30% per year.
  • Full resilience: The single point of failure was eliminated, ensuring high availability through load distribution across multiple replicas.
Picture of Ivan Blanco

Ivan Blanco

Did you find it interesting?

Leave a Reply

Your email address will not be published. Required fields are marked *

Related posts

Limitless scalability: the power of Zabbix Proxy and its automation within the IOMETRICS® Observability ecosystem

AI: To Agent or Not to Agent? That Is Not Always the Question.

Annual customer event: innovation, Autonomous Agents and haute cuisine

Want to know more?