Closing a 92-Day Dense-Compute Evaluation
How two NVIDIA DGX A100 systems completed a bounded 92-day review of power, cooling, monitoring, and operating fit before continued placement.
- Milestone
- Published
- Filed under
- field note
On June 4, 2024, Helixrack closed a 92-day evaluation of two NVIDIA DGX A100 640GB systems that began on March 4. The closeout approved continued placement of both systems within their documented, configuration-specific operating and monitoring boundaries.
The original plan called for a 90-day trial. We kept the evaluation open through the June 4 closeout so the acceptance record, facility observations, maintenance history, monitoring gaps, and customer report could be reconciled before a placement decision was signed. The result applied only to these two systems in their reviewed positions. It was not blanket approval for the DGX A100 product family, another rack position, or a future workload.
Start with a bounded plan
The reviewed systems were 6U rackmount units. NVIDIA’s published DGX A100 specifications list a 6.5 kW maximum system-power envelope. That maximum described the product boundary used for planning; it was not a claim that either system continuously drew 6.5 kW during the evaluation.
Each trial placement used a custom dual 40 A, 208 V A/B power arrangement. The reviewed design allowed either assigned side to support the documented system envelope. That was a configuration-specific trial condition, not the standard power arrangement included with a public Helixrack plan.
Before the evaluation began, the acceptance plan identified both systems, their assigned positions, power inputs, cooling orientation, facility conditions to observe, customer confirmation needed at handoff, and the decision expected at closeout. That scope kept the work focused on fit instead of treating a recognizable model name as sufficient approval.
The review also separated responsibilities. We observed the physical environment and facility-side handoff. The customer retained control of firmware, operating system, applications, credentials, data, and workload-specific acceptance.
Review power and airflow together
Both systems were placed at conditioned-air supply with hot exhaust directed toward the return path. The private evaluation record identified inlet and exhaust observations, room conditions, sensor positions, sampling intervals, and the thresholds that required review.
Those details mattered because rack height and a manufacturer maximum do not describe the complete facility fit. The closeout considered the assigned A and B paths separately, the observed operating states, airflow through each chassis, neighboring equipment, cabling, maintenance clearance, and the response available if a monitored condition moved outside the reviewed boundary.
This summary omits the electrical one-line, rack positions, thresholds, and measured power and temperature values. Those details belong to the two approved placements and would not establish capacity for a different system.
Close against the original questions
At closeout, we compared the observation record with the acceptance plan created before go-live. The review covered each assigned rack position, facility-side power behavior, inlet and exhaust conditions, reachability, maintenance events, monitoring gaps, and customer reports made during the 92-day window.
Each signal describes a different layer. An incoming-power event does not automatically mean the server stopped. A successful reachability check confirms a network path at that moment; it does not establish application health. Customer acceptance answers the questions defined for that configuration and period, rather than creating a general service-level claim.
Against those original questions, the record supported continued placement of both systems. Continued approval depended on preserving each evaluated configuration and staying inside its documented operating and monitoring boundary.
Record the limits with the result
The closing record should identify:
- the exact hardware and configuration evaluated;
- the dates and conditions covered;
- the facility and customer signals used;
- known monitoring gaps or maintenance exclusions;
- observed conditions that stayed within or exceeded the approved boundary; and
- the decision to continue, change, or end the placement, including any conditions attached to that decision.
Without those details, a single percentage cannot tell a future customer whether a different system will fit or how it will behave. This closeout therefore publishes no uptime, application-availability, latency, or performance result. The bounded operating result is narrower: these two placements were approved to continue under the conditions reviewed.
Turn the closeout into an operating decision
The June 4 record closed the evaluation by approving both systems within their reviewed configurations, rack positions, operating conditions, and monitoring limits. It also set a durable rule for later reviews: dense systems begin with their own chassis, power, airflow, weight, cabling, management, and workload information.
That rule protects both sides. The customer receives a decision about the machine being proposed, while we avoid applying an earlier result to hardware that may behave differently.
Begin the next request from zero
Another DGX A100—or any other dense system—may differ in firmware, power supplies, accelerators, fan behavior, airflow, depth, weight, cabling, and workload. It receives its own review rather than inheriting this approval.
The closeout establishes this rule: dense capacity is approved for a documented machine in a documented position. If you are planning one, send the complete configuration and expected operating conditions through Helixrack.com so we can review the fit.
Sources
- NVIDIA announces the DGX A100 640GB configuration November 16, 2020