Deployment and resiliency planning
This guide helps you choose how to deploy ITRS Analytics based on what matters most to your organization, including where the platform is hosted, who operates it, resiliency, and backup capabilities. It also explains the trade-offs associated with each option. Your choice of deployment model directly affects high availability, operational continuity, and your ability to meet uptime and compliance requirements.
This page uses the deployments described in Deployment options:
- ITRS Managed
- Client-managed cluster
- Application on VM (Embedded Cluster)
Note
Please contact your ITRS Account Manager before starting any installation.
Match your operating model Copied
Use the table to match your operating model to a deployment based on who runs ITRS Analytics and whether your organization already operates Kubernetes.
| Who operates ITRS Analytics | Does your organization already operate Kubernetes? | Deployment |
|---|---|---|
| ITRS | Not applicable. ITRS Cloud Operations run the platform. | ITRS Managed |
| Your organization | Yes. Kubernetes is part of your platform strategy. | Client-managed cluster |
| Your organization | No. Kubernetes expertise is limited or unavailable. | Application on VM |
ITRS Managed is used when ITRS should run operations, upgrades, and infrastructure, and data residency can be met by an available AWS region. Your organization self-hosts when data must stay on-premises or in a private cloud that ITRS Managed regional options do not cover, or when procurement, security, or compliance require full infrastructure ownership.
ITRS Managed compared with self-hosted Copied
ITRS Managed is ITRS Cloud Operations running the platform. Client-managed cluster and Application on VM are self-hosted: your organization installs and operates ITRS Analytics. Between those two self-hosted paths, Client-managed cluster is for an existing Kubernetes cluster (for example EKS, GKE, or AKS). Application on VM is for virtual machines, where Embedded Cluster (k0s) provides Kubernetes.
| Feature | ITRS Managed | Self-hosted |
|---|---|---|
| Hosting location | AWS Public Cloud, managed by ITRS | On-premises, private cloud, or public cloud via Client-managed cluster or Application on VM (Embedded Cluster) |
| Platform management | ITRS Cloud Operations teams | Your organization’s operations team |
| Data residency | Client’s choice of AWS Region (costs may vary by region) | Client’s full choice of location and environment |
| Security | TLS, mTLS, Equinix Cloud Connect | TLS, mTLS, service mesh |
| High availability | Deployment spans two availability zones within a single region | Requires an HA Kubernetes design: minimum three controller nodes for Application on VM (Embedded Cluster) HA, workload replicas distributed across failure domains, and a customer-managed load balancer; see resource and hardware requirements |
| Data backup | Daily automated backups | For Client-managed cluster, backup and restore is supported with Velero; your team schedules and runs it |
| Upgrade management | Planned upgrade and client-approved | Client’s responsibility for image mirroring and upgrade execution |
| Application customization | Full identity, role, and access management available with SSO | Full identity, role, and access management available with SSO |
| Cost model | Software costs plus hosting costs (Small, Medium, Large, Extra-Large sizes); optional 2 TB additional storage | Software costs only; infrastructure costs are the client’s responsibility |
With self-hosted deployments, your internal teams are responsible for patching, upgrades, backup management, and maintaining infrastructure resiliency.
Client-managed cluster compared with Application on VM Copied
Client-managed cluster typically uses network-attached persistent volumes. Workloads may be rescheduled to surviving nodes after a failure if sufficient spare capacity exists and the volumes remain accessible over the network.
Application on VM (Embedded Cluster) uses node-local persistent volumes. If a node fails, workloads that depend on its storage cannot be rescheduled to another node. Application-level replication allows ITRS Analytics to remain fully functional when the cluster has sufficient spare capacity, but redundancy and fault tolerance remain reduced until the affected node returns.
| Feature | Client-managed cluster | Application on VM (Embedded Cluster) |
|---|---|---|
| Platform ownership | Customer-managed Kubernetes (EKS, AKS, or GKE) | Kubernetes bundled and managed through ITRS-packaged K0s |
| High availability | Full HA with workload rescheduling across surviving nodes when a node fails, provided sufficient spare capacity exists and network-attached storage remains accessible | Full HA through application-level replication; with sufficient spare capacity, ITRS Analytics remains fully functional after a node failure, but redundancy and fault tolerance are reduced until the node returns because workloads that depend on node-local storage cannot be rescheduled |
| Storage architecture | Persistent volumes are decoupled through storage classes; supports dynamic expansion | Storage is tied to local node disks; data loss risk if a node fails without HA configured |
| Backup and restore | Backup and restore is supported with Velero; your team schedules and runs it | Infrastructure-level backup should be used; Velero does not support the node-local filesystem storage classes used by Embedded Cluster, so use VM snapshots from the hypervisor or storage-level snapshots from the underlying storage platform |
| Load balancing and networking | Supports native cloud and enterprise load balancers with DNS integration, such as AWS NLB, Azure Load Balancer, GCP Load Balancer, or F5 | No built-in load balancer; customers must supply and manage their own, such as HAProxy, keepalived, F5, AWS NLB, Azure Load Balancer, or GCP Load Balancer. A bundled software load balancer is not included because these solutions depend on environment-specific external network cooperation such as ARP/GARP or BGP, which is commonly restricted or unsupported across cloud and many on-premises networks |
| Security | Kubernetes-native security model integrates cleanly with platform operations and customer-controlled controls such as network policies, admission controllers, and pod security standards | Host-level security controls such as antivirus, EDR agents, SSL/TLS inspection, and host firewalls must be validated and exclusions configured before installation; these controls can block image pulls, container runtime activity, or inter-node communication |
| Operational responsibility | Clearly divided across infrastructure, platform, and application teams; cluster issues resolved at the appropriate layer | Cluster issues must be escalated to ITRS because your organization does not have direct access to the bundled K0s Kubernetes layer or its diagnostic tools |
| Maintenance and patching | Integrates with existing customer patching and lifecycle processes for Kubernetes, OS, storage, and networking | Increased coordination risk; patching and upgrades may require downtime and careful change management |
| Disaster recovery | Not built-in; deploy multiple independent ITRS Analytics instances for DR | Not built-in; deploy multiple independent ITRS Analytics instances for DR |
Both self-hosted paths support full high availability. In multi-node deployments, no node should have a round-trip time (RTT) greater than 10 ms to any other node in the Kubernetes cluster.
Important
ITRS Analytics does not provide built-in cross-site or cross-region disaster recovery. For protection against data center or regional failures, you must run multiple independent deployments and implement your own DR strategy (sync, failover, runbooks).
Non-HA configurations Copied
A non-HA single or multi-node Client-managed cluster configuration is used for proof-of-concept and smaller production environments where high availability is not a strict requirement. Proof-of-concept deployments come with no guarantee of high availability for stored data due to their exploratory nature.
A non-HA Application on VM (Embedded Cluster) configuration has additional limitations around data protection:
- Velero-based backup and restore is not supported for node-local filesystem storage classes; use infrastructure-level backups such as hypervisor VM snapshots or storage-platform snapshots instead.
- Risk of complete data loss if a node fails catastrophically.
- Requires complete rebuild if storage is lost.
Warning
Before using Application on VM (Embedded Cluster) in production, plan and validate an alternative backup strategy such as hypervisor VM snapshots or storage-level snapshots. Without a tested backup strategy, a node failure that causes disk loss can result in permanent data loss with no recovery path.
To continue planning:
- For install paths, see Deployment options.
- For backup and restore procedures, see Backup and restore.
- For infrastructure sizing and resource requirements, see ITRS Analytics Sizer.