RWTH High Performance Computing (HPC)
Mehr Informationen zu dem Service finden Sie in unserem Dokumentationsportal.
Nodes unavailable and Jobs stuck
We must report that compute nodes within Claix2023, Claix2025 as well as IH and private systems are currently suffering from random operating system failures that cause the jobs to become unavailable and for Jobs to hang while finishing in a completing state.
We know that this causes the waiting times to increase and that users are unable to stop their hanging jobs.
We are currently working on solving the problem with the uttermost priority.
Until the problem is resolved, we will manually stop users jobs in a stuck state and reset the nodes to a usable state manually.
This also takes time and work, so we ask users to please have patience until it is resolved.
To minimize node downtime, we ask users to try and use single node jobs where possible.
We apologize for this inconvenience and hope to have the issue resolved as soon as possible.
We have identified the issue and are consulting with the vendor.
We have tracked down the responsible bug and received a software update from the vendor. As of now, the new version is rolled out successively on all cluster nodes to fix the issue.
The fixed kernel is now deployed on all nodes. We will monitor the situation and have no reason to believe that jobs will be affected anymore.
ssh in Batchjobs deaktiviert
Due to latent bugs in resource accounting, ssh-ing into nodes where one of your own jobs is running is temporarily disabled.
We are aware of the popularity of this feature and aim to reenable it in a future maintenance.
Slurm privacy rules enforced
We have enforced the Slurm privacy settings and only personal jobs and settings can be seen now. Exceptions cannot be made
https://blog.rwth-aachen.de/itc-changes/en/2026/08/13/slurm-privacy-rules-enforced/
Kürzlich abgelaufene Meldungen
Filesystem outage
Due to a filesystem outage, running jobs may fail and new jobs may not be scheduled.
We are hoping to resolve the issue during working hours on monday.
Jobs still do not run, we'll try to reenable the partitions tomorrow
All Claix 2023 Nodes (were possible) have been released and are up and rurnning.
Claix 2025 Nodes will follow when they are ready.
Zugang zu HPC-Login-Systemen aufgrund einer kritischen Sicherheitslücke gesperrt
Aufgrund einer kritischen Sicherheitslücke im Linux-Kernel wurde der Zugang zu den HPC-Login-Systemen gesperrt.
Sobald Updates verfügbar sind und eingespielt werden konnten, werden wir den Zugang wieder freigeben.
The claix23 dialog nodes (login23-*) are now open for login again.
Teilstörung wurde behoben.
Wartungsarbeiten MySQL Datenbankserver - MFA eingeschränkt
Im angegebenen Zeitraum führt das IT Center Wartungsarbeiten an den MySQL-Datenbanken der Multi-Faktor-Authentifizierung (MFA) durch. Im Verlauf der Wartung werden die Datenbanken für insgesamt etwa 15 Minuten nicht zur Verfügung stehen.
Während dieses Zeitraums sind keine neuen Anmeldungen an Services mit RWTH Single Sign-On möglich. Bereits bestehende Sitzungen bleiben jedoch bestehen und sind von der Wartung nicht betroffen.
Wartung
Due to scheduled maintenance, the entire CLAIX-2023 and 2025 clusters will be unavailable during the specified maintenance window.
We will be performing upgrades to the file systems of CLAIX2025 to upgrade performance and stability upgrades to the filesystem of CLIAX2023. Claix 2023 might be available sooner before the end of the window.
Login nodes will not be available until later in the maintenance window.
Thank you for your understanding.
We have released the Claix2023 nodes but the Claix2025 nodes are still undergoing updates.
Auch die Claix-2025-Systeme sind wieder in Betrieb. Die Wartung ist damit abgeschlossen.
copy23-2 temporarily unavailable
Due to maintenance work and debugging, copy23-2 is unavailable at the moment. Please use copy23-1 in the meantime.
Teilwartung wurde abgeschlossen.