RWTH High Performance Computing (HPC)
Mehr Informationen zu dem Service finden Sie in unserem Dokumentationsportal.
ssh in Batchjobs deaktiviert
Due to latent bugs in resource accounting, ssh-ing into nodes where one of your own jobs is running is temporarily disabled.
We are aware of the popularity of this feature and aim to reenable it in a future maintenance.
With the new Rocky 9.8 updates done, we will next deploy cgroups v2 and re-enable ssh into compute nodes. The timeframe for this is still unfortunately unknown.
Slurm privacy rules enforced
We have enforced the Slurm privacy settings and only personal jobs and settings can be seen now. Exceptions cannot be made
https://blog.rwth-aachen.de/itc-changes/en/2026/08/13/slurm-privacy-rules-enforced/
Kürzlich abgelaufene Meldungen
CLAIX-2025 GPU nodes still stuck
Wir hoffen, das Problem heute beheben zu können.
Es sind noch ein paar wenige Knoten gestört, die Queues sind aber wieder freigegeben
Nodes unavailable and Jobs stuck
We must report that compute nodes within Claix2023, Claix2025 as well as IH and private systems are currently suffering from random operating system failures that cause the jobs to become unavailable and for Jobs to hang while finishing in a completing state.
We know that this causes the waiting times to increase and that users are unable to stop their hanging jobs.
We are currently working on solving the problem with the uttermost priority.
Until the problem is resolved, we will manually stop users jobs in a stuck state and reset the nodes to a usable state manually.
This also takes time and work, so we ask users to please have patience until it is resolved.
To minimize node downtime, we ask users to try and use single node jobs where possible.
We apologize for this inconvenience and hope to have the issue resolved as soon as possible.
We have identified the issue and are consulting with the vendor.
We have tracked down the responsible bug and received a software update from the vendor. As of now, the new version is rolled out successively on all cluster nodes to fix the issue.
The fixed kernel is now deployed on all nodes. We will monitor the situation and have no reason to believe that jobs will be affected anymore.
We have updated the problematic filesystem corrupting the nodes and leaving the jobs hanging. We do not expect any new CG Jobs or Nodes.
Slurm-Update
For urgent reasons, we need to update several components of the Slurm workload management system. This update will entail short periods of unavailability in which no new jobs can be submitted and no information on jobs can be queried from the system. We expect that the update will have no major impact on running jobs.
The update has been successful and to the best of our knowledge no running jobs were negatively affected by it.
We had to update at short notice to mitigate a security vulnerability disclosed today by SchedMD.