Operations

Managed Operations

24×7 operations for mission-critical GPU infrastructure — on our capacity or on yours, run by local teams under a contractual SLA.

Whether the capacity is ours or yours, the same local engineers are on it: continuous on-site operations, RMA coordination and incident response under defined SLAs, with proactive incident management across every active estate.

An engineer in a navy polo working at the front of a live server rack in a data hall aisle.
Support structure

Two tiers, one window

TierRolePosition
L1Data centre operations + RMA engineersStationed on site
L2Network + GPU server engineersSenior engineering
PMProject / service delivery managerSingle contact window
SLA framework

Response by priority, in the contract

PriorityLevelResponse
P1Critical1-hour response
P2High4-hour response
P3StandardNext business day

Continuous monitoring — proactive incident management across all active deployments.

Running today

Active managed services

ID

Indonesia

32× · 32-server GPU estate

Active managed service
SG

Singapore

31× · GPU servers on an InfiniBand fabric

Ongoing Day 2 operations
MY

Malaysia

64× · Rack-scale GPU cluster for a Tier-1 data centre operator

Under monitoring

Services include: L1 DC operations · GPU server management · network monitoring · RMA coordination · SLA incident response.

Incident response

Seven stages, one goal — recovery

  1. 01

    Detection

    Identify and validate

    Identify issues through monitoring, alerts or user reports.

  2. 02

    Assessment

    Understand the impact

    Evaluate severity, scope of impact and affected systems.

  3. 03

    Containment

    Stabilise and limit impact

    Take immediate measures to prevent the impact from spreading.

  4. 04

    Investigation

    Find the root cause

    Investigate the root cause and the path to resolution.

  5. 05

    Recovery

    Return to normal operations

    Implement the fix and restore systems and services.

  6. 06

    Verification

    Ensure systems are stable

    Confirm service stability and monitor full recovery.

  7. 07

    Closure & RCA

    Learn and keep improving

    Document findings and implement improvements.

An engineer at a multi-monitor operations console in a bright data centre control room

Guiding principles: Protect people and systems · Act fast and responsibly · Collaborate effectively · Communicate clearly · Keep improving

AIDC RMA

End-to-end RMA, detection to closure

  1. 01

    Fault detection

    Monitoring or alert triggers, user-identified issues, ticket raised to RCS support.

  2. 02

    Troubleshooting & verification

    Initial troubleshooting, hardware fault verified, affected assets identified.

  3. 03

    RMA request

    RMA raised with the OEM with required information; OEM approval.

  4. 04

    Replacement

    OEM approves the RMA, replacement shipped, defective part returned.

  5. 05

    Install & verify

    Replacement installed, function verified, system confirmed operational.

  6. 06

    Closure

    Ticket updated, resolution documented, RMA closed.

A single server drawn halfway out of its rack on rails in a bright data centre aisle, the empty bay visible beside identical installed units

Covers all in-service assets: GPU servers · CPU servers · network equipment · storage · other infrastructure.

  • All RMA activity follows OEM policy and warranty terms.
  • RMA turnaround depends on the OEM and parts availability.
  • RCS monitoring integrates alerting, tracking and reporting.
Get started

Tell us what you need to run.

Share your workload requirements, market focus, and timeline. Our expert teams across Taiwan, Singapore, Malaysia, and Indonesia will design the optimal compute service—leveraging our existing GPU capacity or building a dedicated infrastructure tailored to your needs.

Talk to our team →