Downtime 20th – 24th of April is over. Services are back in production

All services on Fram and NIRD are now be back in production, except for slurmbrowser and desktop.fram.sigma2.no.

Here is a list of what has been done during the last four days:

  • Firmware upgrade on NIRD in Trondheim and Tromsø
  • Firmware upgrade on NIRD Toolkit
  • Firmware upgrade on Fram storage and Fram nodes, switches m.m
  • Software/OS upgrade on NIRD Trondheim and Tromsø
  • Software/OS upgrade on NIRD Toolkit
  • Software/OS upgrade on Fram nodes

In total, including vendors, ca 15 people were involved in the upgrade.

We thank you for your patience.

tos-project3 on NIRD is read only

Due to underlying hardware issues, tos-project3 filesystem is set to READ-ONLY while we investigate the issue.

These are the projects affected:

NN9999K
NS1002K
NS4704K
NS9001K
NS9012K
NS9014K
NS9033K
NS9054K
NS9063K
NS9066K
NS9114K
NS9191K
NS9320K
NS9404K
NS9518K
NS9602K
NS9615K
NS9641K
NS9672K
NS0000K
NS1004K
NS9000K
NS9003K
NS9013K
NS9021K
NS9035K
NS9060K
NS9064K
NS9081K
NS9133K
NS9305K
NS9357K
NS9478K
NS9560K
NS9603K
NS9616K
NS9655K
NS9999K

Maintenance on NIRD, NIRD Toolkit and Fram , 20th April -24th April

23 April – 18:50 NIRD and the NIRD toolkit services are now back into production

24th April: Fram is back in production.

WARNING: MAINTENANCE IS CURRENTLY ONGOING!

Dear NIRD, NIRD Toolkit, and Fram User,

We will have a four day long scheduled maintenance on NIRD, NIRD Toolkit and Fram starting on the 20th of April, 09:00 AM.

Running HPC jobs and logging in to Saga is NOT affected.
NIRD connectivity, and backup of files, from Saga IS affected

During the maintenance we will:

  • carry out software and firmware updates on all systems

Files stored on NIRD will be unavailable during the time of the maintenance and therefore so will be the services. This will of course affect the NIRD file systems available on Fram and Saga too.

Login services to NIRD, NIRD-toolkit and Fram will be disabled during the maintenance

Please note that backups taken from the Fram and Saga HPC clusters will also be affected and will be unavailable during this period.

Please accept our apologies for the inconvenience this downtime is causing.

Metacenter Operations

NIRD: file system problems

Dear NIRD user,
We have had serious problems with the GPFS file systems this afternoon and had to stop the storage and all the services.

The NIRD storage and the NIRD-toolkit are now back online.
Please notify the metacenter support if you notice any remaining issues.

We are very sorry for the inconvenience.

Update 10:00 24.02.2020 :  We still have problem with Nird mount points on Fram, we are working on the problem, we will keep users posted here.

Update 10:45 24.02.2020:  Problem with Nird mount point on Fram is resolved.

NIRD project file systems mounted on Saga

Dear Saga User,

We have the pleasure to announce that we have now fixed all the technical requirements and mounted NIRD project file systems on Saga login nodes.

You may find your projects in the

/nird/projects/nird

folder.

Please note that to transfer of large amount of files is sluggish and has a big impact on the I/O performance. It is always better to transfer one larger file than many small files.
As an example, transfer of a folder with 70k entries and about 872MB took 18 minutes, while the same files archived into a single 904MB file took 3 seconds.

You can read more about the tar archiving command by reading the manual pages. Type

man tar

in your Saga terminal.

Metacenter Operations

Reorganized NIRD storage

Dear NIRD User,

During the last maintenance we have reorganized the NIRD storage.

Projects have now a so-called primary site which is either Tromsø or Trondheim. Previously we had single primary site, Tromsø. This change had to be introduced to prepare coupling NIRD storage with Saga and the upcoming Betzy HPC clusters.

While we are working on a final, seamless access solution regardless of the primary site for your data, please use the following temporary solution:


To work closest to your data you have to connect to the login nodes located at the primary site of your project:

  • for Tromsø the address is unchanged and is login.nird.sigma2.no
  • for Trondhein the address is login-trd.nird.sigma2.no

To find out the primary site of your project log in on a login node and type:

readlink /projects/NSxxxxK

It will print out a path starting either with /tos-project or /trd-project.
If it starts with “tos” then use login.nird.sigma2.no.
If it starts with “trd” then use login-trd.nird.sigma2.no.

Metacenter Operations

Network outage

Update

  • 2020-01-13 14:54: Problems have been sorted out now and network is functional again.
  • 2020-01-13 14:40: Problems are unfortunately back again. Uninett’s network specialists are working on solving the problem as soon as possible.
  • 2020-01-13 14:22: Network is functional again. Apologies for the inconvenience it has caused.

We are currently experiencing network outage on Saga and some parts of NIRD. The problem is under investigation.

Please check back here for an update on this matter.

Metacenter Operations

NIRD and NIRD Toolkit scheduled maintenance

Update:

  • 2020-01-23 17:30: Services are now progressively restarted.
  • 2020-01-22 21:49: We have detected file system level corruption and to avoid data corruption we had to unmount and rescan all the file systems (about 18PB) on NIRD.
    We are currently working on starting back the services on NIRD Toolkit.
  • 2020-01-22 11:11: Software and firmware is now upgraded on NIRD Toolkit.
    Most of the fileset changes are also carried out. We are currently working on the last bits. Will keep you updated.
  • 2020-01-20 08:58: Maintenance has started. NIRD file systems are unmounted from Fram until maintenance is finished.

Dear NIRD and NIRD Toolkit User,

We will have a three day long scheduled maintenance on NIRD and NIRD Toolkit starting on the 20th of January, 09:00 AM.

During the maintenance we will:

  • carry out software and firmware updates,
  • change geo-locality for some of the projects,
  • replace synchronization mechanisms,
  • depending on part delivery times from disk vendor – expand the storage and quotas.

Files stored on NIRD will be unavailable during the time of the maintenance and therefore so will be the services. This will of course affect the NIRD file systems available on Fram too.

Please note that backups taken from the Fram and Saga HPC clusters will also be affected and will be unavailable during this period.

Please accept out apologies for the inconvenience this downtime is causing.

Metacenter Operations

NIRD crash.

NIRD storage system was crashed and unavailable for short period of time.
Due to this crash, users logged in to NIRD and Fram experienced problemes.
The problem is resolved, NIRD storage system is online now.

Please contact us if you still encounter problems.

Note: The export of NIRD to FRAM does not work currently