Technical architecture
The home page explains what we deliver. This page shows how: the reference topology, the patterns behind HPC, EDA, AI and hybrid multi-cloud estates, the tools we standardise on at each layer, and the terminology in plain language.
Sites on the left, the primary data center in the middle, public clouds on the right. Compliance runs underneath all of it. Every engagement starts from this shape and removes what is not needed.
Engineers work through remote sessions, so design data never lands on a laptop. Offices, remote staff and partner sites all enter through one secure edge with SD-WAN, firewalling and zero-trust access.
A spine-leaf Ethernet fabric carries general traffic; a separate InfiniBand or RoCE fabric carries HPC and AI traffic. Storage is tiered by IO pattern. Identity, licenses and sessions are the middleware that every job depends on.
Clouds are used for what they are best at: burst capacity, GPUs on demand, disaster recovery and SaaS. Private interconnects keep latency predictable; a landing zone keeps every account governed from day one.
Each is a proven topology, adjusted to the measured workload rather than copied from a vendor reference.
For each layer: the tools we standardise on, the terms that come up in design reviews, and the topology pattern we apply.
Layer
Widely used tools
Terminology
Topology and pattern
Where services run when they do not need the whole machine.
VMware vSphere / ESXi, KVM with libvirt, Proxmox VE, Nutanix AHV, OpenStack, Red Hat OpenShift Virtualization (KubeVirt), Docker and Podman for containers.
Type-1 hypervisor, HCI, vCPU oversubscription ratio, live migration, snapshots, SR-IOV, PCIe and GPU passthrough, vGPU, NUMA-aware placement, hugepages.
Bare metal for latency-sensitive EDA, HPC and GPU nodes. Virtualized clusters for identity, license servers, CI, monitoring and web services. HCI for small sites. Containers for stateless services and inference.
How sites, clouds and data centers are joined securely.
Cisco Catalyst SD-WAN (Viptela), Cisco Meraki, Fortinet Secure SD-WAN, VMware VeloCloud, Versa Networks, Palo Alto Prisma SD-WAN.
Underlay and overlay, IPsec tunnels, application-aware routing, path quality (loss, latency, jitter), QoS, hub-and-spoke versus full mesh, SASE, ZTNA, zero-touch provisioning, cloud on-ramp.
Hub at the primary data center or colocation, spokes at offices and partner sites, on-ramps into AWS, Azure and Google Cloud. Dual WAN at every site (fibre plus 5G) with automatic failover and per-application steering.
Where design data, simulation output and training sets live.
NetApp ONTAP, Dell PowerScale, Pure Storage, VAST Data, WEKA, Lustre, BeeGFS, IBM Storage Scale (GPFS), Ceph, MinIO, ZFS and TrueNAS.
NFS v3/v4, SMB, S3 object, parallel file system, NVMe-oF, RDMA, IOPS versus throughput versus latency, metadata performance, tiering (hot, warm, cold), snapshots, replication, erasure coding, RPO and RTO.
Three tiers: scratch (parallel, NVMe, not backed up), project (NAS with snapshots and replication), archive (object, erasure coded). EDA is metadata-heavy, so the project tier is all-flash NAS; AI training sets sit on the parallel tier or object with a local cache.
The CPU fleet, and how it is provisioned and tuned.
Intel Xeon, AMD EPYC, Arm (Ampere, NVIDIA Grace). RHEL, Rocky Linux, AlmaLinux, Ubuntu. Warewulf, xCAT, Foreman, Kickstart and PXE, Ansible, Spack, Environment Modules and Lmod.
Cores versus threads (SMT), base and turbo clocks, NUMA, memory bandwidth and channels, hugepages, CPU pinning, BIOS and power profiles, stateless (diskless) nodes, image-based provisioning, kernel tuning.
EDA tools are often single-thread bound: high-frequency nodes with large memory per core. HPC favours high core counts and memory bandwidth. Nodes are grouped into homogeneous pools, one per scheduler queue, provisioned from a single golden image.
GPUs and other accelerators, and the fabric between them.
NVIDIA H100, H200 and B200 (HGX and DGX systems), AMD Instinct MI300, Intel Gaudi, FPGAs from AMD (Xilinx) and Intel (Altera), dataflow accelerators. CUDA, ROCm, oneAPI, TensorRT, NCCL and RCCL.
NVLink and NVSwitch, PCIe Gen5, MIG (multi-instance GPU), time-slicing, GPUDirect RDMA and GPUDirect Storage, HBM, FP8 and BF16 precision, tensor cores, TDP and rack power density, liquid cooling.
Eight-GPU nodes with NVSwitch inside the node; a rail-optimised InfiniBand NDR or 400GbE RoCE fabric between nodes. GPU sharing (MIG, time-slicing) for development and inference. Bare-metal pods with Kubernetes or Slurm on top. Power and cooling designed before the servers are chosen.
What decides which job runs where, and when.
Slurm, IBM Spectrum LSF, Altair Accelerator and PBS Pro, OpenPBS. Kubernetes (OpenShift, Rancher, RKE2, EKS, AKS, GKE), Helm, Argo CD, Flux, KubeVirt, Ray, Kubeflow, HashiCorp Nomad.
Fair-share scheduling, priority and preemption, backfill, gang scheduling, queues and partitions, resource reservations, license-aware scheduling, autoscaling, node pools, GPU operator, job arrays.
Slurm, LSF or Altair for batch HPC and EDA farms; Kubernetes for services, inference and MLOps. Cloud bursting through AWS ParallelCluster, Azure CycleCloud or Google Cluster Toolkit with the same queue names, so users do not notice where a job ran.
Identity, licenses, sessions and the glue between tools.
FlexNet Publisher (FlexLM), RLM, Altair Monitor, OpenLM. FreeIPA / Red Hat IdM, Active Directory, Keycloak, Okta, Microsoft Entra ID. NICE DCV, NoMachine, ThinLinc, VNC. GitLab, Artifactory, HashiCorp Vault.
LDAP, Kerberos, SSSD, SSO, SAML, OIDC, MFA, RBAC, license checkout and queueing, session broker, API gateway, message queue (Kafka, RabbitMQ), secrets management, PKI.
Identity is the control plane: one directory, federated to every cloud. License servers run highly available and network-close to the compute they serve. Session brokers sit in front of remote visualization so engineers get a desktop inside the data center from anywhere.
Data center fabric, campus and wireless, and the low-latency fabric.
Cisco Nexus and Catalyst, Arista, Juniper. NVIDIA Quantum (InfiniBand) and Spectrum (Ethernet). Palo Alto Networks, Fortinet. Cisco and Meraki, Aruba for Wi-Fi 6E/7. Infoblox, NetBox, WireGuard and IPsec.
Spine-leaf (Clos), VXLAN/EVPN, BGP, OSPF, MLAG and LACP, 25/100/400GbE, InfiniBand HDR and NDR, RoCEv2, PFC and ECN (lossless Ethernet), east-west versus north-south traffic, microsegmentation, out-of-band management, 802.1X.
Two-tier spine-leaf for the data center; a separate non-blocking fabric for HPC and AI; a dedicated out-of-band management network. Campus with Wi-Fi 6E/7 and 802.1X, and zero-trust access for remote engineers instead of a flat VPN.
Policy, evidence and protection across people, systems and models.
Wazuh, Splunk and Elastic Security (SIEM). CrowdStrike and SentinelOne (EDR). Tenable Nessus, Qualys, OpenVAS. CyberArk (PAM). OpenSCAP and CIS-CAT. AI-DLP and prompt-security gateways in front of external LLMs.
ISO 27001, SOC 2, NIST CSF, CIS Benchmarks, DPDP Act (India), GDPR, export-control awareness, least privilege, MFA, audit trail, policy as code, DLP, AI usage policy, prompt injection, shadow AI.
Identity and MFA everywhere; networks segmented per project or customer; central logging into a SIEM; an AI gateway that redacts sensitive data before it reaches external models. Evidence is collected continuously by the platform, not assembled the week before an audit.
The terms that appear in our proposals and design reviews, each in one sentence.
Spine-leaf
A two-tier switch design where every leaf connects to every spine, giving predictable latency and simple scale-out.
VXLAN / EVPN
Overlay networking that stretches networks across the fabric, with BGP distributing the control information.
RDMA, RoCE, InfiniBand
Memory-to-memory transfers that bypass the CPU; the basis of every low-latency HPC and AI fabric.
NVMe-oF
NVMe flash storage accessed over the network at close to local speed.
Parallel file system
Storage that stripes data across many servers so thousands of nodes can read and write at once (Lustre, GPFS, WEKA).
Tiered storage
Placing data on hot, warm or cold tiers by how it is used, to balance performance against cost.
Scheduler and fair-share
The batch system that queues jobs and allocates resources by policy (Slurm, LSF, Altair), so no one team starves another.
License server
The service that meters EDA tool licenses (FlexNet, RLM); the most common hidden bottleneck in a design farm.
Session broker and remote visualization
Engineers work on desktops inside the data center; only pixels cross the network (NICE DCV, NoMachine).
MIG
Multi-instance GPU: splitting one GPU into isolated slices so several small jobs share it safely.
NCCL
The collective-communication library that lets GPUs across many nodes train as one system.
Kubernetes
Container orchestration; the control plane for services, inference and MLOps pipelines.
Infrastructure as code
Environments defined in versioned files (Terraform, Ansible) and rebuilt on demand, identically, every time.
GitOps
The Git repository is the source of truth; every change ships through a reviewed pull request (Argo CD, Flux).
Landing zone
A pre-governed cloud foundation (accounts, IAM, networking, logging) that every workload lands on.
Hybrid cloud and multi-cloud
Hybrid joins on-premises with a public cloud; multi-cloud uses more than one provider. Most estates are both.
Cloud bursting
Sending overflow jobs to cloud capacity when the local farm is full, then releasing it when the peak passes.
Data gravity
Large datasets pull compute toward themselves; it is cheaper to move the job than the data.
SD-WAN overlay
Encrypted tunnels over any underlay (fibre, broadband, 5G), steered per application by policy.
SASE and ZTNA
Access to applications based on identity and device health rather than network location; no flat VPN.
Identity federation (SSO, SAML, OIDC)
One login, trusted across on-premises tools and every cloud console, with one MFA policy.
Observability
Metrics, logs and traces in one place (Prometheus, Grafana, Zabbix, Elastic) so problems are seen before users report them.
RPO and RTO
How much data you can afford to lose and how long you can afford to be down; the two numbers behind every backup and DR design.
FinOps
Treating cloud spend as an engineering metric: tagged, forecast and right-sized every month.
AI-DLP
Scanning prompts and uploads so confidential data never reaches an external model unprotected.
HCI
Hyperconverged infrastructure: compute and storage in one scale-out cluster, common for small sites and services.
Share your workload profile and constraints. We will return a topology, a bill of materials and an operating plan.