Cooling issues in Fram server room

Update:

  • 11:15  Cooling distribution units are functional again and computes started back once again.
  • 10:12  Cooling units failed once again and computes were automatically switched off. We are looking into the problem.
  • 09:22 Cooling is functional again and Fram computes are started back now and machine shall shortly be fully operational.

——————————————————————————————-

We had troubles with one of the cooling units in the Fram server room today around 06:30.
Safety mechanisms switched off biggest part of the Fram compute nodes.

Thank you for your understanding!
Metacenter Operations

Fram: scheduled downtime on the 28th of August

UPDATE

2018-08-28 18:02 Fram is up and jobs running again.

We will have a one day scheduled downtime on Fram on the 28th of August starting from 08:00 AM.

Jobs not being able to finish before the maintenance window, will be left pending in the queue with a Reason “ReqNodeNotAvail” and will be started when the maintenance is over.

We will keep you updated via OpsLog/Twitter.

Thank you for your consideration!
Metacenter Operations

Issues with job completion – FIXED

Update : 14:39 26-07-18 The issue with Fram file system is now fixed and jobs should run as normal.

We are experiencing some problems at the moment and this is most likely a file system issue. We are trying our best to bring the services back to normal, however as most of the experts are on holiday this may take longer than usual. Please check back here for updates.

Cooling system issues – some jobs crashed

A sudden high pressure in the cooling system around 13 o’clock, has taken one of the cooling units down. Starting it back affected the other unit as well.
This triggered a safety stop for some of the computing nodes, leading to premature crash for some of the running jobs.

Affected jobs has been re-queued.

Apologies for the inconvenience it has created.
Metacenter Operations