QuickMDSim
All posts

Up to 64 cores, and a clearer CPU or GPU choice

QuickMDSim 8 min read

A LAMMPS job on QuickMDSim can ask for up to 64 cores. The machine exists only while the job runs. Extra cores always cost more per second. They do not always finish sooner. We timed that, on two systems, plus one NVIDIA L4.

Two choices, one credit wallet

The Run control is still one picker. CPU is 1, 2, 4, 8, 16, 32, or 64 cores. GPU is one NVIDIA L4, on Pro. GPU is not “64 cores with a graphics card.” It is one GPU, and it burns 12 CPU-units per wall-second. A 64-core CPU job burns 64.

Eight cores is a different machine than four. The charts below do not connect those two classes into one line.

What we measured

Rates are timesteps per second while LAMMPS was actually stepping, not time spent waiting for a machine. The same input ran at every size.

Two charts. Left: 256,000-atom Lennard-Jones timesteps per second from 8 to 64 cores, with the L4 as a dashed line far above. Right: 32,000-atom PPPM, almost flat from 8 to 64 cores, L4 only slightly above the best CPU.
Figure 1. Loop rate. Left: short-range LJ, 256,000 atoms, 1,000 steps. Right: charged LJ plus PPPM, 32,000 atoms, 100 steps. Dashed line is one L4. CPU points are 8–64 cores.
Bar charts of timesteps per credit-second. For the large LJ run the L4 bar is much taller than 8, 16, 32, or 64 cores. For PPPM the 8-core bar is the tallest and 64 cores is the shortest.
Figure 2. Same runs, divided by the credit multiplier (cores, or 12 for the L4). Higher means more steps per credit-second.

256,000 atoms, short-range LJ

FCC Lennard-Jones liquid, reduced density 0.8442, lj/cut 2.5, NVE, timestep 0.005. 40×40×40 cells, 256,000 atoms. Neighbor skin 0.3, rebuild every 20 steps.

  • 8 cores: 45.1 steps/s. Pair was 70% of the loop.
  • 16 cores: 50.0 steps/s. Almost no gain. Communication rose to 41%.
  • 32 cores: 110.7 steps/s. Best CPU. 2.5× the 8-core rate, not 4×.
  • 64 cores: 87.6 steps/s. Slower than 32. Communication was 76% of the loop.
  • L4: 366.8 steps/s. 3.3× the best CPU, 8.1× the 8-core rate.

On credits the L4 is the cheap one here: 30.6 steps per credit-second, against 5.6 on 8 cores and 1.4 on 64. Sixty-four cores was the most expensive way to run this script, and not the fastest.

Every size, including the L4, printed the same final thermo line (step 1000, T = 0.545, PE = −6.097). Same trajectory, not just a run that finished.

32,000 atoms, LJ plus PPPM

Same density, charges of +1 and −1 alternating (net charge 0), lj/cut/coul/long 2.5, pppm 1.0e-4. LAMMPS built a 120×120×120 grid. The long-range part was 95% or more of the loop at every size. The FFT on this GPU is not a fast one. That is the result.

  • 8 cores: 8.75 steps/s.
  • 16 cores: 8.79 steps/s.
  • 32 cores: 10.5 steps/s. Best CPU, and only 20% above 8 cores.
  • 64 cores: 8.73 steps/s. Same as 8 cores, eight times the credits.
  • L4: 10.6 steps/s. Ties the best CPU. Slower per credit than 8 cores (0.89 vs 1.09 steps per credit-second).

The 8-core and 16-core runs match the L4 energy and temperature. On 32 and 64 cores the starting potential shifted by under 2%. That is the long-range grid split across cores, not a different system. The rate comparison still holds: the FFT did not get faster as cores were added on this one machine.

2,048 atoms, the small case

Bar chart for 2,048 Lennard-Jones atoms. 1 core near 600 steps per second, 2 cores near 1,100, 4 cores back near 500, and the L4 near 8,100.
Figure 3. Same LJ liquid, 2,048 atoms, 2,000 steps. The L4 is a separate machine from the 1–4 core runs.

One core: 602 steps/s, pair 83% of the loop. Two cores: 1,124 steps/s. Four cores: 513 steps/s, slower than one core, 77% of the loop in communication. The L4 did 8,123 steps/s, 13.5× the one-core rate, and slightly more steps per credit (677 vs 602). Four cores were the worst buy on this script.

The 256,000-atom box, melting

The pictures are a separate run of the same 256,000-atom LJ box, not a timing run. A hot start (T = 2.5) on the FCC lattice, then 800 steps of NVE. Potential energy went from −6.77 to −5.02. Mean displacement from the initial site was 0 at step 0 and 1.23σ at step 800. 89% of atoms had moved more than 0.6σ. The lattice is gone.

Two panels of a thin slab through 256,000 Lennard-Jones atoms. Left: an ordered lattice at step 0. Right: the same slab at step 800, disordered.
Figure 4. A 3.4σ slab through the box, viewed from the side. Left: step 0, still on lattice sites. Right: step 800. Slate is still near the initial site. Amber has moved. About 16,000 atoms in the slab at step 0.
Figure 5. Cutaway of all 256,000 atoms, same run, 21 frames, played forward and back. The return to the lattice is the player, not the physics. Color is displacement from the step-0 site.

How to decide

From these three scripts:

  • Short-range LJ at 256,000 atoms: 32 cores beat 8 and beat 64. The L4 beat every CPU size on time and on credits.
  • PPPM at 32,000 atoms, this accuracy, this grid: more cores did not pay. Eight cores beat the L4 on credits. The L4 did not beat 32 cores on time.
  • LJ at 2,048 atoms: two cores helped. Four cores did not. The L4 was much faster and not more expensive per step than one core.

The picker states the burn rate before you press Run, and it will not start a size your balance cannot cover for one minute. A run that starts and then fails still uses credits for the time the cores were reserved. A job that never gets a machine does not.

Who can pick what

  • Free: 1 core.
  • Basic: 1–16 cores.
  • Pro: 1–64 cores, and GPU.

One wall-hour on 64 cores is 64 CPU-hours. That is most of a Pro month. On the 256,000-atom LJ script, that hour would also have been slower than 32 cores.

This is a single machine, not a multi-node cluster. We do not promise a speedup. The timings above are the reason.

LAMMPS 22 Jul 2025. Cite Thompson et al., Comp. Phys. Comm. 271, 108171 (2022).