GPU-native explicit finite-element solver

More physics.
Less waiting.

Put GPU computation to work on your next engineering question.

MinuteSim brings GPU-resident explicit simulation to medium-to-large structural workloads. Explore more designs. Build toward richer physics data for AI.

Actual S-rail forming result, colored by effective plastic strain at the shell top fibre, with the original 0 to 0.40 color scale.
ACTUAL MinuteSim RESULTA sheet becomes a structure.
EXPERIMENTAL
Adaptive shell refinement
01

GPU-residentExplicit time integration

02

Shells · solids · beamsSupported configurations

03

FP32 & FP64Single and double precision builds

02 / The value of shorter runtimes

Large models.
A different timescale.

Make room for the next design iteration. This solid-compression scenario shows how the CPU comparison basis changes the runtime ratio.

EXPERIMENTAL · ILLUSTRATIVE SCENARIO
590×

Illustrative CPU-to-GPU runtime ratio

Supplied timing assumptions · one configuration
SOLVER B / 1 CPU THREAD22.6 hoursEstimated from a short timing window
MinuteSim / ONE NVIDIA L402.3 minutes137.8 s projected at normal GPU clocks
1,886,592 solid elements0.5 s physical analysis duration119,821 MinuteSim time increments*

FP32 GPU / extended-single CPU configuration on separate EPYC 9274F-class servers. Time increments and output differ. Final-field equivalence is not established for this timing case. *Increment count comes from the corresponding completed GPU run.

Read the assumptions and observed runtimes

A scenario for one model, not a product-wide speed guarantee. The observed 8-thread CPU and GPU full-run times are available alongside the projections.

EXPERIMENTAL · ILLUSTRATIVE SCENARIOS

Runtime scenarios across model sizes.

One NVIDIA L40 versus Solver B. Compare the same problem families with 8 CPU threads and 1 CPU thread.

Sheet forming

A sheet stretches over a rounded punch.

Deformed sheet over a rounded punch, colored by shell-top effective plastic strain, with the original 0 to 0.25 color scale.

Representative result · 19,881 elements

Mixed-element bending

Shell and solid specimens bend under tooling.

Four bent specimens made with triangle shells, quad shells, tetrahedra and hexahedra, with tooling and the original effective-plastic-strain color scale.

Contact illustration · timing series: no contact

Solid compression

A rigid hemisphere presses into a solid block.

Quarter view of a rigid hemisphere compressing a solid block, colored by effective plastic strain with the original 0 to 0.50 scale.

Representative result · 384,000 elements

Match each model title’s color to its curve below. Images show representative results; the curves cover multiple model sizes.

8 CPU thread runtime scenarios. At the largest plotted sizes, CPU-to-GPU ratios are 132 times for sheet forming, 46 times for mixed-element bending and 92.4 times for solid compression. Both axes are logarithmic.
8 CPU threads / one L40View full size ↗
1 CPU thread runtime scenarios. At the largest plotted sizes, CPU-to-GPU ratios are 517 times for sheet forming, 314 times for mixed-element bending and 590 times for solid compression. Both axes are logarithmic.
1 CPU thread / one L40View full size ↗

X: element count. Y: CPU time / GPU time. Both panels use the same logarithmic axes.

Illustrative scenarios using supplied timing assumptions, including recorded timings and estimates. Endpoint labels refer to the largest plotted model in each family. Bending timing uses prescribed motion without contact; the gallery shows separate contact examples. Comparison conditions and data ↗

03 / Physics data for AI

1,000 simulations
for AI training.

Representative physics data is part of the foundation for surrogate models. Faster simulation can change the cost of exploring that data space.

PLANNED APPLICATION

Explore the compute-time scenario.

Multiply a recorded single-run time by a chosen number of sequential simulations.

1002,000

MeshGraphNets used 1,000 training trajectories per dataset. It is an example of dataset scale, not a universal requirement. Read the study ↗

Sheet forming

19,881 shells
CPU · 8 threads
79 days
One L40 GPU
18.5 hours

Per run: CPU 6,839.55 s · GPU 66.676 s

Mixed-element bending

64,800 elements · contact
CPU · 8 threads
16.5 days
One L40 GPU
2.6 days

Per run: CPU 1,429.374 s · GPU 226.817 s

0 days50 days100 days

Both charts share a linear scale. CPU: Solver B on 8 EPYC 9274F threads. GPU: one NVIDIA L40. Single-precision configurations.

Illustrative arithmetic, not measured sustained throughput or a completed dataset. Sequential runs on one CPU allocation or one GPU; concurrent CPU jobs would change the comparison. Data preparation and AI model training time are excluded. Scenario details ↗

The hardware opportunity

GPU hardware.
A wider horizon for physics data.

Selected hardware generations show how compute and memory bandwidth have developed. This creates room to explore larger models and more physics data for AI.

GPUCPUL40 / benchmark GPU
FP32 compute for selected desktop CPUs and GPUs from 2008 to 2025, on an absolute TFLOPS axis. Endpoints: RTX 5090 at 104.9 TFLOPS and Ryzen 9 9950X at 4.4 TFLOPS. The separate L40 reference is 90.5 TFLOPS.
FP32 compute / selected desktop productsView full size ↗
Memory bandwidth for selected data-center GPUs and server CPU sockets, 2007 to 2025, on an absolute GB per second axis. Endpoints: B300 at up to 8,000 GB/s and Xeon 6980P at 844.8 GB/s. The L40 reference is 864 GB/s.
Memory bandwidth / GPU or CPU socketView full size ↗

The MinuteSim comparisons on this page use NVIDIA L40.

SUPPORTED hardware specifications. Absolute linear axes. GPU values follow published FP32 conventions; CPU compute values are theoretical at base clocks. Growth labels use each series’ own starting value. B300 bandwidth is up to 8,000 GB/s for the selected configuration. The L40 marker uses 2023 OVX system availability. These specifications do not predict solver speed or compare matched price and power. Products, conventions and sources ↗

Historical source data: Karl Rupp, CPU, GPU and MIC Hardware Characteristics over Time, CC BY 4.0, and Epoch AI, Data on machine learning hardware (CC BY 4.0), plus selected vendor specifications. Figures redrawn with absolute axes and an added L40 reference; graph extracts from the v11 presentation.

04 / The solver today

A broader set of
engineering building blocks.

SUPPORTED · DEVELOPMENT CAPABILITIES

Supported combinations vary by element, material, contact and constraint. Availability in the documented beta is narrower.

01

Elements

  • Hex8 solids
  • Tet4 and Tet10 solids
  • Wedge6 solids
  • Quad4 and Tri3 shells
  • Beam elements
Element scope ↗
02

Materials & failure

  • Linear elasticity
  • Tabulated and bilinear plasticity
  • Strain-rate scaling
  • Barlat ’89 sheet plasticity
  • Neo-Hookean rubber
  • Rigid tooling
  • Quad4 shell erosion subset
Material scope ↗
03

Contact & rigid walls

  • Node-to-surface contact
  • Sheet-forming contact
  • Coulomb contact friction
  • Planar rigid walls
  • Moving spherical walls
  • Friction on planar walls
Contact scope ↗
04

Loads & constraints

  • Prescribed motion
  • Nodal and rigid-body forces
  • Gravity and initial velocity
  • Fixed degrees of freedom
  • MPC · linear constraints
  • RBE2 · rigid coupling
  • RBE3 · weighted interpolation
Constraint subsets ↗
Selected keyword inputMinuteSim / NVIDIA GPUResults for ParaViewEXPERIMENTAL · adaptive shell refinement

Let’s explore the fit

Your model.
The next conversation.

Discuss a representative workload, the results that matter, and a technical evaluation under agreed conditions.

FOR CAE TECHNOLOGY & PRODUCT TEAMS

  • Representative model evaluation
  • Technology collaboration
  • Strategic partnership discussions
[email protected]