Planned maintenance on Fram on 28.11.2019

Update:

  • 2019-11-28 15:00: Fram is opened for production and scheduled jobs were resumed.
  • 2019-11-28 14:33: For more information about enforced file system quota accounting, please see this announcement.
  • 2019-11-28 13:26: Lustre quota indexes are fixed now. Running some additional cleanup and tests.
  • 2019-11-28 08:03: Maintenance has started.

Dear Fram User,

The slave quota indexes on the /cluster file system got broken and as a result of it the quota accounting has become unpredictable. For the same reason some of you may have even experienced job crashes.

To fix this, we will have a two days planned downtime starting from 08:00AM on the 28th of November.

Fram jobs which can not finish by the 28th of November, are queued up and will not start until the maintenance is finished.

Thank you for your consideration!

Metacenter Operations

Planned maintenance on Fram on 16.10.2019

Update:

  • 2019-10-18 14:36 We are ready with the reinstallation, configuration checks, QA and tests. Access to the machine has been reopened and queued jobs are running again.
  • 2019-10-18 06:12 Reinstallation of compute nodes is much slower then anticipated and thus re-opening of the machine is delayed. We do our best to finish the maintenance as soon as possible. In parallel we are conducting tests and benchmarks.
    Will keep you updated.
  • 2019-10-17 08:25 File system servers and infrastructure switches were patched yesterday.
    We are proceeding now with the upgrade of the service and the login nodes.
  • 2019-10-16 08:07 Maintenance has started.

Dear Fram User,

We will have a two days planned downtime starting from 08:00AM on the 16th of October for maintenance on the storage and the file system.

During this time we will, together with the vendor, upgrade the storage firmwares, upgrade the software on the /cluster file system servers and upgrade the operating system on Fram.

This upgrade is necessary to fix the frequent issues with the metadata servers and enhance stability and security of the system.

Fram jobs which can not finish by the 16th of October, are queued up and will not start until the maintenance is finished.

Thank you for your consideration!

Metacenter Operations

NIRD and Service Platform downtime – 26.06.2019

  • 2019-07-02 08:00: All Service Platform services resumed. It might be that some of the services are not properly working and need to be restarted after the maintenance. If you experience any problem with your service, please do not hesitate to contact us asap.
  • 2019-06-27 19:58: NIRD filesystems are mounted back to Fram.
  • 2019-06-27 19:46: NIRD login nodes are started back now, you may login and access your files stored on NIRD.
    Remaining Service Platform services will be started tomorrow morning.
  • 2019-06-27 09:54: We are starting back and testing the file system now.
  • 2019-06-26 22:08: All hardware replacements are done now and the storage system is monitored for any signs of instability. Starting back of the filesystem is planned for tomorrow morning. We will keep you updated.
  • 2019-06-26 14:05: Vendor is meticulously checking each NIRD storage component and decided to replace main controller chassis.
    In the mean time we are applying firmware updates on the Service Platform to improve stability and security.
  • 2019-06-26 08:15: Maintenance has started.

Dear NIRD and Service Platform User,

We have a planned downtime on the 26th of June, Wednesday next week, to replace some defective hardware. Systems will taken offline starting from 08:00AM.

Engineer from storage vendor will assist us from the very first hour.

We expect the maintenance to finish in one day.
Will keep you updated here.

Metacenter Operations

Scheduled downtime – NIRD storage expansion – 2nd of April

Update:

  • 2019-04-03 13:25: NIRD and the service platform are back into production.
  • 2019-04-03 10:59: Maintenance work has finished. We are proceeding in starting back the filesystems and services.
  • 2019-04-03 08:22: Disk expansion and rebalancing is finished. HW checks are currently ongoing and shall finish in a couple of hours. Will keep you posted.
  • 2019-04-02 09:55: NIRD filesystems are unmounted from Fram and replicated data is available read-only trough login-trd.nird.sigma2.no
  • 2019-04-02 08:06: Maintenance work has started.

Dear NIRD User,

NIRD and the Service Platform will be under maintenance to expand the disk capacity in Tromsø.

The operations for storage expansion and disk pool rebalancing will start on the 2nd of April at 8:00 am CET and will last for maximum 2 days. During the maintenance, the services running on the NIRD Service Platform and on the NIRD Toolkit will not be available.

During the downtime we plan to make project data mirrored to Trondheim available in read-only mode trough a specially built login node. This solution will be first tested with real load during this downtime, thus we might encounter some technical difficulties.
That being said, to access the remote, mirrored data, please login to login-trd.nird.sigma2.no.

We apologise for the inconvenience.
Metacenter Operations

Scheduled downtime on the 12th of February

Update:

  • 2019-02-26 08:57: Fram parts have arrived and were installed yesterday. The vendor will start rebuilding the disk pools on Fram today.
  • 2019-02-22 09:10: Access to NIRD is reopened. We will proceed in taking up the Service Platform during today. Some disk pools are still under rebuild and should be finished in few hours. So until then, you might encounter slight performance loss.
  • 2019-02-21 15:40: Most of disk rebuild on NIRD are finished. NIRD file system is started and we will proceed with opening access to NIRD as soon as remaining network issues are sorted out on the backbone.
    Due to logistic issues, Fram parts are held back at the customs. The vendor sent a second batch of parts through another logistic company and ETA is Monday morning.
    We will post a message when access to NIRD is re-opened.
  • 2019-02-19 14:10: Disk rebuilds on NIRD have reached 40%. Current ETA for NIRD is Thursday morning.
    We are still waiting for the Fram parts to arrive to Tromsø.
  • 2019-02-18 16:36: Some disk pools must be rebuilt once again for NIRD, thus delaying the opening once more. We are terribly sorry for this. Will continue updating the log as soon as new information is available.
  • 2019-02-18 10:10:
    NIRD storage is stabilized now and the vendor will do a new attempt of taking the system back online during today. At this stage it is still uncertain when Fram can be put back into production.
  • 2019-02-15 11:18: We are still experiencing problems with the storage system on Fram. Disks begun to mass-fail once again after the system seemed to be stable during the night. We are depending on the vendor to resolve these issues and we are working closely with them.
    Based on the new instability we can not give an estimate for when the system will be ready for general use again. This is an unfortunate situation and we understand the impact on you, and thus we try all possible solutions to keep your data safe and bring up the system as soon as possible.The OpsLog will be updated with new information when the status of the situation changes.
  • 2019-02-14 13:17: Due to missing parts, and the size of the storage, disk recovery is progressing slowly ahead on approximately 50% reduced performance. Current ETA are:
    • Fram: 15.02.2019
    • NIRD: 19.02.2019
    • Service Plattform: 19.02.2019
  • 2019-02-13 19:07: Communication with the missing storage enclosures were re-established and disk pools are rebuilding at this time. Unfortunately we can not reopen machines until disk pools are stabilized. We will have a new round of checks and risk analysis tomorrow morning. Will keep you updated here.
  • 2019-02-13 11:33: Some of the parts arrived to the datacenter and we are working with the vendor on replacing and pathing the firmware on Fram. More details to follow as we know more.
  • 2019-02-12 15:38: NIRD Tromsø and Fram storages have each one disk enclosure which failed. We are waiting for replacement parts to arrive. After replacement we will have to rebuild disk pools before re-opening machines for production.Current estimate is tomorrow evening. Will keep you updated.
  • 2019-02-12 12:36: Firmware upgrade on NIRD is finished. We are proceeding to start back NIRD services. Will keep you posted.
  • 2019-02-12 08:17: Maintenance has started.
  • 2019-02-11 13:20: Due to the disk problems accellerating during the weekend, we have now changed the maintenance stop reservation so no new jobs will start until the maintenance is done.  Already running jobs will not be affected, but no new jobs will start.  This has been done to reduce the risk of data loss.

We need to have a scheduled downtime on a relatively short notice in order to upgrade the firmware on both Fram and NIRD (including NIRD Toolkit) storages.
This is a critical and mandatory update which will increase stability, performance and reliability of our systems.

The downtime is expected to last no more than a working day.

Fram jobs which can not finish by the 12th of February, are queued up and will not start until the maintenance is finished.

Thank you for your understanding!
Metacenter Operations

Slurm Upgrade on Fram

Update: The upgrade is now done. All seems to have gone well, but it can be a good idea to check your jobs.

We will upgrade the queue system (Slurm) on Fram at 11:00 today, from version 17.11 to 18.08. The upgrade is expected to take 5-10 minutes. During that time, queue system commands (squeue, sbatch, etc.) will not work, but running jobs should not be affected.

Fram, NIRD and Service Platform scheduled downtime starting on 21st November 2018

UPDATE:

  • 2018-11-23 09:45:
    • We plan to re-open access to Fram during today. Will keep you updated.
    • We are currently running benchmarks and tests on the upgraded system.
    • Cooling units are stress-tested now to pinpoint any outstanding issues.
    • OpenMPI is upgraded now.
  • 2018-11-22 22:48:
    • Lustre servers are upgraded on Fram and several tests, benchmarks were run to fine tune parameters.
    • First step of the CDU maintenance is carried out now.
  • 2018-11-22 11:33:
    • NIRD Service Platform is up again.
    • Access to NIRD is reopened. Please note that we have now four login nodes and SSH fingerprint is changed.
  • 2018-11-21 18:41:
    • Needed hardware replacement for NIRD was carried out and all firmware upgrades are finalized on both Tromsø and Trondheim site storage systems.
    • NIRD file systems are started back now and we plan to reopen access tomorrow before noon.
    • Firmware is updated now on the Fram storage system.
    • Several other updates, including the OpenFabrics stack and Lustre, were done in parallel.
  • 2018-11-21 08:00: Maintenance has started.

Dear Fram, NIRD and Service Platform User,

On the 21st and 22nd of November we will have a scheduled maintenance on Fram, NIRD and Service Platform.

This will be a comprehensive maintenance on the national HPC and research data infrastructure, ongoing on multiple levels and sites. Due to it’s complexity and amount of work involved, some parts of the infrastructure might require downtime extension for the 23rd of November, too.

The work will include, but will not be limited to:

– firmware upgrades on disks, enclosures, chassis, etc.

– operative system upgrades

– queue system upgrade

– file system upgrades

– kernel upgrades

– upgrade of OpenMPI

– upgrades on the OpenFabrics stack

– maintenance on the cooling system units

 

Our aim is to enhance the stability and security of the infrastructure, eliminate bugs and enhance performance, while having the shortest downtime possible.

We understand that system unavailability has big impact on your daily work and such we try to bring back our systems functional as soon as possible.

 

Thank you for your consideration!

Metacenter Operations

Fram: scheduled downtime on the 28th of August

UPDATE

2018-08-28 18:02 Fram is up and jobs running again.

We will have a one day scheduled downtime on Fram on the 28th of August starting from 08:00 AM.

Jobs not being able to finish before the maintenance window, will be left pending in the queue with a Reason “ReqNodeNotAvail” and will be started when the maintenance is over.

We will keep you updated via OpsLog/Twitter.

Thank you for your consideration!
Metacenter Operations