The same loop can support certificate renewal, password rotation, image compliance, drift remediation, patch planning, or service provisioning. The controls should become stricter as the potential blast radius increases. A low-risk catalog deployment can be highly automated. A fleet-wide lifecycle change still needs staged execution, prechecks, change gates, and recovery planning.

Fleet Scope

For each proposed domain, record:

Instance Scope

A practical example is capacity expansion. Demand increases in the AI domain. Telemetry identifies sustained accelerator pressure. Policy confirms that the threshold and budget conditions have been met. An approved workflow adds or assigns capacity. Validation confirms host health, network readiness, cluster state, licensing, and workload placement. The outcome is recorded for cost and capacity planning.

Domain Scope

Test site loss, management-service degradation, network isolation, identity failure, certificate expiration, capacity exhaustion, failed upgrades, and recovery sequencing. The topology should explain how operators detect the problem, contain it, continue essential services, and restore control.

Tenant and Application Scope

AI infrastructure often creates the strongest case for a dedicated domain because accelerator hardware, high-speed networking, data access, driver compatibility, capacity economics, and security requirements can diverge sharply from general-purpose virtualization.

An AI/GPU domain can establish a controlled landing zone for deep learning virtual machines, GPU-enabled Kubernetes clusters, inference services, model-development environments, or private AI services. The boundary can also make expensive capacity visible, protect accelerator availability, and isolate changes that depend on specialized firmware, drivers, device profiles, and networking.

Where the Archipelago Model Breaks Down

A useful workload-domain contract should answer:

Too Many Islands

The operating model should also name the bridges explicitly: identity, DNS, NTP, certificate authorities, backup targets, logging, monitoring, repositories, automation endpoints, service registries, external networks, and support escalation. A shared service without an owner is an unplanned common failure domain.

One Giant Island

Use lifecycle, hardware, trust, availability, capacity, ownership, cost, and dependency density to determine whether a separate domain is justified. Record the decision and the conditions that would cause it to be revisited.

Ungoverned Bridges

Get Paul Bryant’s practical guides to enterprise AI, hybrid platforms, and day-2 operations by email. New articles as they publish. Unsubscribe anytime.

False Autonomy

First, central visibility should not be confused with identical service levels. The platform can observe several domains through a common operations layer while each domain retains different performance, lifecycle, and recovery objectives.

Capacity Islanding

A workable ownership model separates responsibilities by scope:

Recovery Inside the Same Failure Domain

The management domain is the special-purpose foundation for the VCF instance. It supports the components and relationships required to operate the environment. That makes its availability, capacity reservation, backup, certificate lifecycle, identity integration, monitoring, and recovery posture materially different from a general application landing zone.

A Practical Workload-Domain Design Sequence

The design question is not, “What kind of workload is this?” The better question is, “What must be operated differently for this workload to meet its service objectives?”

Establish Platform Invariants

The key lesson is that a workload domain is not simply a collection of clusters with a convenient label. It is a deliberate boundary around infrastructure characteristics and operational responsibility.

Build Workload Profiles

A Kubernetes-oriented domain can provide a governed foundation for VMware vSphere Kubernetes Service clusters and the supporting platform capabilities around them. The design focus shifts from individual virtual machines to namespaces, cluster lifecycle, container networking, registries, policy, observability, and platform-team service levels.

Score the Boundary Signals

Done well, the result is not a collection of infrastructure silos. It is a private cloud that can support traditional applications, Kubernetes platforms, AI services, and recovery operations through one governed architecture while preserving the boundaries that make production operations manageable.

Map Shared Services and Traffic

Dedicated hardware protects service levels, but it can also strand expensive capacity. AI and GPU domains need utilization targets, quota policy, reservation rules, and a process for rebalancing or expanding resources.

Define the Domain Charter

The uploaded image gives us a better starting point. It shows several illuminated cities rising from a shared infrastructure ocean. Each city has a different purpose. One represents traditional enterprise systems, another modern application platforms, another AI, and another security or recovery. A protected central city coordinates the environment while telemetry panels watch capacity, health, threats, and demand.

  • Purpose and allowed workload classes
  • Accountable platform owner
  • Hardware, storage, and network profile
  • Security and identity boundary
  • Lifecycle and maintenance cadence
  • Capacity reservation and growth model
  • Availability, RTO, and RPO objectives
  • Backup and recovery dependencies
  • Required shared services
  • Automation and self-service interfaces
  • Monitoring, cost, and compliance evidence
  • Exit, consolidation, or decommission criteria

Validate the Topology Against Failure

The recovery city in the image needs the strongest terminology guardrail. A separate domain can reserve recovery capacity or isolate protection components, but disaster recovery is not achieved merely by creating another workload domain inside the same failure boundary.

Automate Only After the Contract Is Clear

Instance scope includes the management domain, core VCF relationships, instance lifecycle, platform certificates, backups, and the dependency chain required to operate the local domains. This is the level where architects should document what happens when management services are degraded but workloads continue running.

Operational Implications for VCF 9.1

The metaphor is useful, but it can encourage weak designs if taken too literally.

Define the services and standards that should remain consistent across the private cloud: identity, naming, time, certificate trust, logging, monitoring, configuration evidence, backup policy, security baselines, and lifecycle governance.


Domain scope includes cluster configuration, host and storage profiles, network connectivity, capacity reservations, lifecycle sequencing, workload eligibility, and domain-specific monitoring. The domain owner should know what can change independently and what still depends on fleet or instance services.


TL;DR

VCF 9.1 strengthens the management model around centralized lifecycle, fleet visibility, management services, observability, and API-driven infrastructure. Those capabilities make the archipelago easier to operate, but they do not eliminate the need for architecture discipline.

A separate enterprise VM domain is justified when these workloads should be insulated from faster-moving platform services or specialized hardware pools. It is less useful when the environment is small and the proposed separation does not change lifecycle, risk, performance, or ownership.

However, a GPU domain is not a substitute for AI governance. Model access, data classification, registry controls, secrets, prompt and output handling, observability, tenant quotas, and cost attribution still require explicit ownership above the infrastructure layer.

Continue reading

External References

The image is useful because it shows separation and connection at the same time. The cities are distinct, but they are not isolated. Bridges carry traffic and services between them. The ocean is shared. The weather affects the whole environment. Central dashboards provide visibility across the system.

VCF 9.1 Multitenant Solar System: Designing Tenant Worlds Around Shared Platform Services
Design multitenancy around shared platform services and explicit tenant contracts. Connect identity, quotas, network isolation, lifecycle ownership, and recovery requirements before deciding where domain boundaries belong.
Read the article →

Similar Posts