Downtime Betzy, 6th December 08:00 until 9th December 08:00

There will be a scheduled downtime for Betzy lasting three days starting on Monday 6th December at 08:00. Downtime will last until Thursday 9th, 08:00.

During the downtime we will conduct:

  • Full upgrade of the Lustre filesystem (both servers and clients)
  • Full upgrade of the infiniband firmware
  • Full upgrade of the Mellanox infiniband drivers
  • minor updates to other parts of the system (Slurm, configs, etc)

Please be aware that this does also affect the storage services recently moved from NIRD to Betzy.

We apologize for the inconvenience

[FINISHED] Short 1 hour downtime on Fram, 24th November 12:00

[Update, 2021-11-24 14:15] Now the NIRD mounts are working again.

[Update, 2021-11-24 13:30] We are back in production and jobs are running som normal again. We are missing the NIRD mounts on two of the login nodes, but are working on fixing that.

[Update, 2021-11-24 12:00] The maintenance has started now.

There will be a short 1 hour downtime for Fram on 24th November, starting at 12:00.

During downtime we will update the firmware on interconnect infiniband switches

[SOLVED] Betzy Downtime 7th June 15:00-20:00

[UPDATE, 2021-06-08 08:00] Betzy is now up and in production again.

[UPDATE] Unfortunately, the downtime is taking longer than anticipated, and will not be finished tonight. We plan on getting Betzy up again at around 08:00 tomorrow morning.

Campusservice at NTNU will conduct maintenance on the High Voltage circuits for Non-redundant power on 7th of June 2021, between 15:00 and 20:00. All compute nodes and login nodes will be shut down during this time, and no jobs will be running during this period. Submitted jobs estimated to run into the downtime reservation will be held in queue.

New compute nodes on Saga

Today, Saga has been extended with 120 new compute nodes, increasing the total number of CPUs on the cluster from 9824 to 16064.

The new nodes have been added to the normal partition. They are identical to the old compute nodes in the partition, except that they have 52 CPU cores instead of 40.

We hope this extension will reduce the wait time for normal jobs on Saga.

Fram: compute nodes are down

Dear Fram users,

We have problem with Fram compute nodes, there are about 870 nodes is down due to unknown reason, we are working on the issue, and will keep you updated.

Update 2020-12-22, 20:05: Most of the compute nodes have now been brought back online. There are still a few nodes that needs more checking before being made available for jobs.

Update 2020-12-22, 18:04: The cooling system has been stable for the last hour after making some adjustments together with the vendor. We are slowly bringing up the nodes.

Update 2020-12-22, 16:01: In order to keep the cooling as stable as possible, we have decided to take down all high memory nodes. This way we can keep some of the normal compute nodes up for the time being. We are also working together with the vendor to make adjustments on the cooling system to ensure continued stability.

We are very sorry about the inconvenience.

Update 2020-12-22, 13:41: We have identified the cause to be the cooling system and are working on mitigating the issues. Most of the compute nodes must remain down while doing so, unfortunately.

Update 2020-12-24 10:30: Compute nodes shutdown again due to electrical problems in machine room, problem has been resolved according to machine room service department, we are working to take up all nodes.

Update 2020-12-24 12:10: Most of the compute nodes on Fram is back online.

[RESOLVED] Saga downtime. 7th December-11th December. Adding 4PB storage.

As previously announced, Saga will be down in the coming week, from 7th December 08:00 until 11th December 16:00.

The downtime is allocated for expanding the storage. When we come back we will have ca 4 Petabyte in addition to the already existing 1 PetaByte.

Update: Saga is back online and running jobs again. The new storage is not online yet, but all the hardware has been mounted.

NIRD-TOS maintenance is active

Dear NIRD and NIRD Toolkit user: NIRD-TOS is currently down and will remain unavailable until Wednesday 12:00. We are replacing all cables during the next coupe of days.

Note that NIRD-home is also not available during that time.

All remote mounts on Fram, Saga and Betzy using NIRD-TOS will be unavailable until downtime is over

Downtime for NIRD-TOS (including toolkit) and Fram 2nd November- 6th November

We will have downtime the following week to try again to replace all internal cables in NIRD-TOS and Fram storage systems.

NIRD-TOS (Including the toolkit) will be down from 08:00 Monday 2nd November to wednesday 4th 12:00

Fram will be down from Wednesday 4th 08:00 until Friday 6th 12:00

There is still a chance that the downtime will not happen, but proper notification will be given in the opslog. Unfortunately the current situation with Covid-19 makes it difficult to make detailed plans.

We apologize for any inconvenience.