Performance / Across model sizes
See the time
behind the speedup.
Compare CPU and GPU runtimes for the same problem family as model size grows. Three workloads. One NVIDIA L40.
Runtime scenarios from the 6 October 2026 comparison tables. Recorded timings and projection methods →
01 / Runtime comparison
Sheet forming
A rounded punch stretches a sheet through an 80 mm stroke. Tool faces are refined with the sheet mesh.
Benchmark setup
- GPU · MinuteSim
- Quad4 · fully integrated assumed strain
- CPU · Solver B
- Quad4 · full integration
Both: 5 through-thickness integration points; 1 mm sheet.
- Contact
- Frictionless node-to-surface penalty contact between the sheet and rigid punch / die.
- Tool & motion
- Rigid rounded punch: 80 mm downward in 0.08 s (1 m/s). The die and sheet-edge constraints remain fixed.

Representative result · 19,881 sheet elements
- Solver B · 1 CPU thread
- 28.21 h
- MinuteSim · one L40
- 3.28 min
1 CPU thread · Runtime
Swipe the chart to see all model sizes →
View all 4 model sizes and times Seconds · CPU 1 & 8 threads
| Sheet elements | MinuteSim / L40 | Solver B / 8 threads | Solver B / 1 thread |
|---|---|---|---|
| 4,900 | 21.693 | 1,010.0 | 3,298.6 |
| 10,000 | 36.455 | 2,612.8 | 9,805.2 |
| 19,881 | 66.676 | 6,839.6 | 28,366.5 |
| 50,176 | 196.600 | 25,928.0 | 101,563.4 |
02 / Runtime comparison
Mixed-element bending
Replicated beam specimens combine four solid and shell element configurations in roughly equal proportions. The timing series uses prescribed displacement without contact.
Benchmark setup
- GPU · MinuteSim
- Hex8 · F-bar (2 × 2 × 2)
- Tet4 · one-point
- Quad4 · fully integrated assumed strain
- Tri3 shells (C0 triangle)
- CPU · Solver B
- Hex8 · full integration (2 × 2 × 2)
- Tet4 · one-point
- Quad4 · full integration
- Tri3 · 3-node shell
Both: 5 through-thickness integration points for shells. Each element family contributes roughly one quarter of the total.
- Contact
- No contact in the timing series. Loading and supports are imposed directly on nodes.
- Tool & motion
- Center top-surface nodes: 3 mm downward over 0.04 s with a smooth displacement ramp. Support span: 192 mm. No moving punch.

Contact illustration · timing series has no contact
- Solver B · 1 CPU thread
- 44.64 h
- MinuteSim · one L40
- 8.54 min
1 CPU thread · Runtime
Swipe the chart to see all model sizes →
View all 6 model sizes and times Seconds · CPU 1 & 8 threads
| Total elements | MinuteSim / L40 | Solver B / 8 threads | Solver B / 1 thread |
|---|---|---|---|
| 9,600 | 14.481 | 166.2 | 940.1 |
| 20,960 | 19.615 | 293.4 | 1,871.0 |
| 50,760 | 36.521 | 1,024.4 | 6,736.4 |
| 100,800 | 57.054 | 1,985.2 | 13,201.3 |
| 203,040 | 183.300 | 7,488.6 | 52,015.4 |
| 494,080 | 512.400 | 23,393.4 | 160,695.1 |
03 / Runtime comparison
Solid compression
A rigid hemisphere compresses a quarter-model solid block through a 500 mm stroke over 0.5 s of physical time.
Benchmark setup
- GPU · MinuteSim
- Tet4 · one-point
- CPU · Solver B
- Tet4 · one-point
Both: 1 integration point per tetrahedral solid element.
- Contact
- Frictionless node-to-surface penalty contact between the solid block and a meshed rigid hemisphere.
- Tool & motion
- Rigid hemispherical punch: 500 mm downward in 0.5 s at a constant 1 m/s.

Representative result · 384,000 solid elements
- Solver B · 1 CPU thread
- 22.58 h
- MinuteSim · one L40
- 2.30 min
1 CPU thread · Runtime
Swipe the chart to see all model sizes →
View all 6 model sizes and times Seconds · CPU 1 & 8 threads
| Solid elements | MinuteSim / L40 | Solver B / 8 threads | Solver B / 1 thread |
|---|---|---|---|
| 82,944 | 18.195 | 520.3 | 2,299.6 |
| 162,000 | 21.040 | 958.5 | 4,430.5 |
| 384,000 | 31.206 | 2,306.1 | 16,051.9 |
| 750,000 | 54.069 | 4,623.2 | 32,427.1 |
| 998,250 | 71.855 | 6,164.6 | 41,676.4 |
| 1,886,592 | 137.800 | 12,736.2 | 81,303.3 |
Reading these comparisons
Time to compute.
Within a defined scope.
Hardware & timing basis
MinuteSim uses one NVIDIA L40 with FP32. Solver B uses 1 or 8 CPU threads on AMD EPYC 9274F-class hardware. The tables combine recorded runtimes with supplied full-run projections. Both CPU selections use the same GPU series.
Compare within each family
Each family has its own loading and element configuration. Element count is a model-size index, not an isolated measure of mesh resolution or hardware scaling. The bending series changes both mesh size and the number of specimen groups. The result images are representative; video duration is not calculation time.
Full comparison details
Thread estimates, normal-clock projections, output settings and the recorded large-compression run are explained on the Evidence page.
Numerical accuracy
These timing comparisons do not establish final-field equivalence. Explore separate accuracy comparisons for the numerical results and their reference conditions.


