Monitoring - All but one hugemem nodes are up, and multiple users have had jobs start and finish as expected. Coldfront should now be operational. DMF is up and looks ok.
OnDemand is available but testing has shown degraded performance. We are continuing to work with our vendors to health check tape drives media.
Please note that REANNZ are unable to test systems where we don’t have access. Please don’t hesitate to contact support with any issue discovered.
We will resume providing updates on Monday.
Jul 31, 2026 - 16:56 NZST
OnDemand is available but testing has shown degraded performance. We are continuing to work with our vendors to health check tape drives media.
Please note that REANNZ are unable to test systems where we don’t have access. Please don’t hesitate to contact support with any issue discovered.
We will resume providing updates on Monday.
Jul 31, 2026 - 16:56 NZST
Update - We are looking into gpu and hugemem nodes, which continue to be unavailable at present.
Service restoration of Coldfront is ongoing but progressing well.
HPE engineers are currently helping investigating the health of DMF and tape media services.
We plan to provide our next update around 17:00.
Jul 31, 2026 - 15:10 NZST
Service restoration of Coldfront is ongoing but progressing well.
HPE engineers are currently helping investigating the health of DMF and tape media services.
We plan to provide our next update around 17:00.
Jul 31, 2026 - 15:10 NZST
Update - We have brought OnDemand, login and all compute nodes back online and are currently monitoring their health status. Jobs can be submitted. GPFS/SMB access via S: and M: drives is available. All jobs that were running either failed or have been cancelled.
Service restoration efforts to Nix, GQuery and Coldfront are ongoing.
We expect to provide our next update around 15:00.
Jul 31, 2026 - 12:57 NZST
Service restoration efforts to Nix, GQuery and Coldfront are ongoing.
We expect to provide our next update around 15:00.
Jul 31, 2026 - 12:57 NZST
Update - Access to eRI via ssh and OnDemand is still closed as we test the compute nodes. When the GPFS nodes are mounted, files will be available via Windows drive access.
We are bringing compute nodes back online but they will not be accessible until we have completed health checks on them. Nix, GQuery, Peaks and related VMs will all be brought online as part of that process.
Service restoration efforts are ongoing. We do not yet have enough information to provide an accurate recovery estimate, but we will continue to share regular updates. Unless the situation materially changes, we expect to provide our next update around 13:00.
Jul 31, 2026 - 09:58 NZST
We are bringing compute nodes back online but they will not be accessible until we have completed health checks on them. Nix, GQuery, Peaks and related VMs will all be brought online as part of that process.
Service restoration efforts are ongoing. We do not yet have enough information to provide an accurate recovery estimate, but we will continue to share regular updates. Unless the situation materially changes, we expect to provide our next update around 13:00.
Jul 31, 2026 - 09:58 NZST
Update - Recovery of login services is progressing well, however this work is ongoing. The compute nodes are not online yet, batch jobs are not running.
We expect to provide an update at approximately 10:00am tomorrow.
Jul 30, 2026 - 17:06 NZST
We expect to provide an update at approximately 10:00am tomorrow.
Jul 30, 2026 - 17:06 NZST
Update - HPC storage services appear to be in good health, although final checks are ongoing.
In parallel, we are working to get login services up and running however, at this stage we do not expect to have compute nodes online today.
Freezer and tape-based storage is expected to remain offline for longer, while we perform drive and media consistency checks. We will provide further information on this as we know more.
We plan to provide another update at approximately 17:00 today.
Jul 30, 2026 - 14:59 NZST
In parallel, we are working to get login services up and running however, at this stage we do not expect to have compute nodes online today.
Freezer and tape-based storage is expected to remain offline for longer, while we perform drive and media consistency checks. We will provide further information on this as we know more.
We plan to provide another update at approximately 17:00 today.
Jul 30, 2026 - 14:59 NZST
Update - We have powered on essential hardware to begin bringing systems online. We are actively working on ramping up services starting with HPC storage and validating the status of storage. We are bringing up the minimum number of virtual instances to manage the rest of the fleet once storage is online.
We are keeping a very close eye on the environmental metrics at the data centre to prevent any additional fluctuations that might impact our tape media.
We plan to provide another update at approximately 15:00 today.
Jul 30, 2026 - 12:56 NZST
We are keeping a very close eye on the environmental metrics at the data centre to prevent any additional fluctuations that might impact our tape media.
We plan to provide another update at approximately 15:00 today.
Jul 30, 2026 - 12:56 NZST
Update - Tamaki Data Centre (TDC) has confirmed that the cooling system is back online and they are performing environmental checks etc.
We need to do a slow and controlled resumption of business as usual and perform health checks on the systems as we go.
We plan to provide another update about the progress and status by 13:00 today.
Jul 30, 2026 - 10:04 NZST
We need to do a slow and controlled resumption of business as usual and perform health checks on the systems as we go.
We plan to provide another update about the progress and status by 13:00 today.
Jul 30, 2026 - 10:04 NZST
Update - The current status is that we have the two working chillers at the data centre - TDC and a team is working on restoring the full chiller capacity.
Jul 30, 2026 - 10:02 NZST
Jul 30, 2026 - 10:02 NZST
Identified - Tamaki Data Centre has identified issues affecting two of the site's three chillers, resulting in reduced cooling capacity.
Teams are actively investigating the root cause and monitoring site temperatures and cooling performance. Our priority is the safe restoration of cooling services and maintaining the integrity of the platform.
Once temperatures have stabilised and cooling capacity has been validated, we will commence our recovery and startup plan. Further updates will be provided as the investigation progresses, this will likely be from tomorrow morning
Jul 29, 2026 - 17:22 NZST
Teams are actively investigating the root cause and monitoring site temperatures and cooling performance. Our priority is the safe restoration of cooling services and maintaining the integrity of the platform.
Once temperatures have stabilised and cooling capacity has been validated, we will commence our recovery and startup plan. Further updates will be provided as the investigation progresses, this will likely be from tomorrow morning
Jul 29, 2026 - 17:22 NZST
Update - We are continuing to investigate this issue.
Jul 29, 2026 - 17:02 NZST
Jul 29, 2026 - 17:02 NZST
Update - Cooling at the Tamaki Data Centre has not yet been restored. All services and hardware remains down. We currently don't have any ETA.
Jul 29, 2026 - 16:24 NZST
Jul 29, 2026 - 16:24 NZST
Update - Contractors are onsite at the TDC but have not yet confirmed the chiller issue
Jul 29, 2026 - 15:21 NZST
Jul 29, 2026 - 15:21 NZST
Update - The data centre has chiller issues - Tamaki Data Centre (TDC) and all nodes are in danger of overheating. Consequently all nodes are being shutdown to protect the hardware
Jul 29, 2026 - 14:59 NZST
Jul 29, 2026 - 14:59 NZST
Investigating - We are currently investigating this issue.
Jul 29, 2026 - 14:57 NZST
Jul 29, 2026 - 14:57 NZST
About This Site
Bioeconomy Science Institute - AgResearch eRI status
Identity Broker Service
Operational
Managed Storage Service
Degraded Performance
General Flexi HPC Platform
Operational
Network connectivity
Operational
Compute cluster
Degraded Performance
Login nodes
Degraded Performance
Operational
Degraded Performance
Partial Outage
Major Outage
Maintenance
Major outage
Partial outage
No downtime recorded on this day.
No data exists for this day.
had a major outage.
had a partial outage.
