sofia @ VUB-HPC#
sofia is the 4th VSC Tier-1 cluster, following muk (hosted by HPC-UGent, 2012-2016), BrENIAC (hosted by HPC-Leuven, 2016-2022) and Hortense (hosted by HPC-Ugent, 2021-2027).
sofia is in production since July 7th 2026 and is hosted by Vrije Universiteit Brussel.
Getting help#
Any questions, comments or problems related to Tier-1 sofia can be addressed to the central support for VSC Tier-1 services:
Questions regarding Tier-1 Hortense should be send to: compute@vscentrum.be.
Hardware details#
The sofia cluster consists of the following compute partitions:
Slurm partition |
|
|
|
|
|---|---|---|---|---|
Nodes |
56 |
16 |
22 |
4 |
GPUs per node |
8x NVIDIA H200 (Hopper) |
2x NVIDIA RTX 5000 Ada (Ada) |
||
GPU memory |
140 GB |
32 GB |
||
CPUs per node |
2x 192-core AMD EPYC 9965
|
2x 96-core AMD EPYC 9655
|
2x 96-core AMD EPYC 9654
|
2x 96-core AMD EPYC 9655
|
CPU memory |
742 GB
|
1493 GB
|
1493 GB
|
1493 GB
|
Local disk |
1.4 TB SDD |
1.4 TB SDD |
1.4 TB SDD |
1.4 TB SDD |
Interconnect |
200 Gbps
|
400 Gbps
|
800 Gbps
|
400 Gbps
|
Shared infrastructure:
storage: 4.3 PiB shared scratch storage, IBM Storage Scale System
interconnect: NVIDIA Quantum-2 InfiniBand, fat tree network with 2:1 blocking
Getting access#
The sofia VSC Tier-1 cluster can only be accessed by people with an active Tier-1 compute project . Everyone is welcome to request a starting grant or collaborative grant to prepare for their Tier-1 compute project application.
Web portal#
The Tier-1 cluster sofia has its own OnDemand Web Portal. Users with an active project in sofia can access it at sofia OnDemand.
More information about the usage of the web portal is available at Access: Web Portal.
Terminal interface#
You can use SSH to connect to the terminal interface of the Tier-1 cluster sofia with your VSC account.
login.sofia.vub.be
There are two ways to authenticate with SSH to sofia, with an SSH key in your VSC account page or with Multi-factor Authentication (MFA). Different restrictions apply to each:
- SSH certificates with MFA (any public network)
Set up your SSH connection to connect to sofia with your VSC ID and a SSH certificate via MFA as described in Connecting with an SSH agent. You can connect to sofia from any network with this method of authentication.
- SSH keys in VSC Account Page (Flemish university networks)
Configure your SSH client as you have already done for other VSC Tier-2 clusters to connect to sofia with your VSC ID and your SSH key of choice in your VSC account page. This method of authentication is restricted to the Flemish university networks. This means that you need to connect to sofia from within your university premises, use your university VPN or connect from another VSC Tier-2 compute cluster.
Login nodes#
There are two login nodes in sofia: login01 and login02.
Upon login you will be assigned to either of these login nodes. If you need to
access a specific login node (for example because you have a screen or
tmux session running there), you can jump between login nodes with the
commands ssh login01 or ssh login02.
Warning
Compute intensive tasks are not allowed on the login nodes. Please use an
interactive job, preferentially on the zen5_vis partition, either via
salloc -p zen5_vis or through the Web portal.
SSH Host keys#
The first time you log in to the sofia login nodes, a fingerprint of the host key will be shown. Before confirming the connection, please verify the correctness of the host key to ensure you are connecting to the correct system.
The fingerprint that will be shown depends on the version and configuration of your SSH client, check that the one shown corresponds to one of the following:
for ECDSA host key:
78:2f:bd:30:95:7f:ba:eb:c4:bf:29:71:fd:1f:2b:b7(MD5)Trw7Gwirs3k9QdH2n2fB9NGrNKPwG8mVK3Yfy8VPc2Y(SHA256)
for ED25519 host key:
bf:4d:d3:55:ce:72:cd:dd:a7:bf:44:62:97:5f:22:2d(MD5)IPZeanj7AVbBYFyc1KBFfd+hMxyDbIGyOghn2uQkuww(SHA256)
for RSA host key:
5e:2e:21:11:c6:8e:1e:1f:c9:7c:d2:38:c7:f6:f9:35(MD5)vS9fAoDO52dWCXS/obPS5u95irhJiKo1PfV4dyJl2Mg(SHA256)
Access and retention policy#
The resources of a Tier-1 compute project are always allocated for a given period of time. Once the project expires, users of the project will not be able to use any more resources granted within that project.
Attention
Data retention of project data is 60 days.
Use of resources stops at the end of the day of the last day of the project. Users of the project will still be able to access their data for 60 days after that date. The following details the changes at each stage once a project ends:
- End date of the project:
SSH login to sofia
OnDemand Web portal of sofia
write data to sofia project directory
access data in sofia project directory via Globus
access data in sofia home directory via Globus
- 60 days after end of project:
all project data in sofia deleted
Managing project members#
Only users who are members of a Tier-1 compute project can access its resources on sofia. The project manager can add other users to the project via the VSC hub.
Log in to the VSC hub with your VSC account.
Open your project and go to its Team tab.
Click Add and select either Member, if the user has already logged in to the VSC hub before, or Invite by email, if the user has never logged in there yet.
Warning
Gaining access to the bsofia_proj_202x_yyy user group through
account.vscentrum.be is not
sufficient to access sofia. You must instead be added as a member of
the project itself in the VSC hub.
Note
Users with a single word Full name (Gecos) in account.vscentrum.be will not be able to properly use the VSC hub and might see No association when attempting to join a project. To resolve this, go to account.vscentrum.be, click Edit account, and change your Full name (Gecos) to be at least two words (e.g. your first name followed by your last name). Then log out and back in again to VSC hub.
Consult your project resource usage#
You can consult your project’s resource usage directly in the VSC hub.
Open your project in the VSC hub and go to its Resources tab to see the compute resources allocated to your project and how much of them have been used.
See also
More information on requesting access, rules and regulations can be found in https://www.vscentrum.be/compute.
Storage#
The Tier-1 cluster sofia has 4.3 PiB of very fast storage. This is a shared storage available on all login and compute nodes of the cluster. It is used to provide scratch storage for jobs (i.e. project directories), as well as user’s home directories and it also holds the installations of scientific software.
Globus on sofia#
The sofia scratch storage can be accessed via the Globus data sharing platform. You can find the link to the endpoints of sofia in Globus in the link below.
Note
Remember to back up your project data in sofia. Data will be deleted 60 days after the project has expired. See our retention policy.
Home directory#
The user’s $HOME directory in sofia is located on its own scratch file
system and is distinct from the user’s $VSC_HOME found in other Tier-2
clusters. The quota for $HOME is 50 GB and 256.000 files.
Users will have a default account setup upon their first login to
sofia. If you want copy any configuration files or customizations (e.g.
.bashrc or any other dot files) from your Tier-2 cluster, you can do so
through Globus.
The size of the $HOME is larger so it can be used to install custom
software or keep tools across projects.
One advantage of this setup is that sofia remains accessible even if the Tier-2 infrastructure on the user’s home institution is down.
VSC_DATA#
Your $VSC_DATA storage, which is hosted in your home institute, is not
directly accessible from sofia.
Users can access their $VSC_DATA storage with Globus
and transfer any needed data in there to their project directory in the scratch
storage of sofia.
Node-local scratch#
Each node in the cluster, including the login nodes, provides local storage for temporary data that is not shared or visible to other users.
Temporary storage on local node hard drive:
$VSC_SCRATCH_NODE,$TMPDIR,/tmpand/var/tmpTemporary storage on local node memory (RAM):
/dev/shm
Note
Data placed in any of these temporary storage locations will be deleted at the end of the active session. For compute nodes this is at the end of your job and for login nodes whenever you log out of the cluster.
Job submission#
The Tier-1 cluster sofia uses the Slurm job scheduler. Only Slurm-native commands are supported for managing your jobs.
Users must specify one of the available partitions when submitting jobs.
Loading a cluster module is not required.
Job environment#
In sofia, both batch and interactive jobs start in a clean session environment. This
differs from the default Slurm behavior (--export=ALL). This approach
provides the following advantages:
reproducibility: your jobs run consistently regardless of your current shell state
architecture alignment: software modules will always be loaded for the hardware architectures in use, ensuring access to the correct and optimized software binaries
Propagating specific environment variables to your job can be done with the
--export=<environment_variables> option in Slurm. See The job environment
for more information.
If your workflow requires your full shell environment to be propagated, please refer to the VUB-HPC documentation on how to copy your full shell environment into your job.
Job memory#
The CPU memory allocated to Slurm jobs scales linearly with the number of allocated CPU cores. See Hardware details for the memory per core available in each partition.
All jobs must use the default memory allocation. Memory overrides --mem,
--mem-per-cpu, and --mem-per-gpu are not allowed and will cause a job
submission error.
This policy avoids situations where CPU cores are available but cannot be allocated because insufficient memory remains available on the node. It also helps keep benchmark results representative of production runs by ensuring that all jobs follow the same linear memory-per-core allocation.
GPU jobs#
In the zen4_h200 partition, Slurm jobs are allocated a fixed ratio of
24 CPU cores per GPU (1/8 of the CPU cores on a node). Job requests
that do not follow this ratio will be rejected. The allocated CPU memory
will scale at the same ratio.
We recommend the following #SBATCH directives to request GPU resources.
The example below requests 2 GPUs on a single node, with 1 task allocated per
GPU (2 tasks in total) and 24 CPU cores per task (48 cores in total):
#SBATCH --partition=zen4_h200
#SBATCH --nodes=1
#SBATCH --gpus-per-node=2
#SBATCH --ntasks-per-gpu=1
#SBATCH --cpus-per-task=24
This policy avoids situations where GPUs are available but cannot be allocated because insufficient CPU cores available on the node. It also helps keep benchmark results representative of production runs by ensuring that all jobs use the same core-to-GPU ratio.