How to Run Java GPU Workloads in Docker

An old GTX 1060, a fresh Docker install, and ten million Monte Carlo paths: from GPU passthrough to Java option pricing with Greeks.

dockergpucudacontainersnablatensor

I'll use a real machine: a small desktop with a GeForce GTX 1060, 6 GB of VRAM, and an Intel i5-7500T. No rented datacenter card, no imaginary terminal output. The finish line is ten million Monte Carlo paths, an option price and five first-order Greeks, computed by NablaTensor inside Docker, then checked against the same computation on the CPU. An old gaming GPU still gets to have a productive afternoon.

Watch the GPU say hello from the host and from Docker, then get to work. Follow live nvidia-smi telemetry and an ASCII scoreboard built from the measured results. A separate 100-million-path load run makes GPU activity visible; the CPU/CUDA comparison uses ten million paths on each backend. Commands are scripted for pacing, and all executions and readings come from real runs. Long idle waits are shortened in playback.

The three pieces you need

PieceLives whereIts job
NVIDIA driverHostTalks to the physical GPU through the host kernel
NVIDIA Container ToolkitHostExposes the requested devices and driver libraries inside Docker
Application and CUDA librariesImageRuns your code; here, Java plus NVRTC, CUDA's runtime compiler

The image shares the host's kernel. Installing a kernel driver in a Dockerfile would put it on the wrong side of that arrangement. NVIDIA's toolkit supplies the bridge; its Docker configuration documentation explains device selection and which driver libraries get mounted.

Here is the tested combination:

ComponentVersion or hardware
Host OSUbuntu 24.04.5 LTS, x86-64
CPUIntel Core i5-7500T, four cores
GPUNVIDIA GeForce GTX 1060, 6 GB, Pascal
NVIDIA driver580.178.04
Docker Engine29.1.3, Ubuntu docker.io package
NVIDIA Container Toolkit1.20.1
Image CUDA release12.9.1
NVRTC packagecuda-nvrtc-12-9, 12.9.86-1
JavaEclipse Temurin 25.0.1+8, from the Maven build image
NablaTensorRevision 26006a665b074d4e231c8dfa405467afa90bc595, CUDA backend, FP32

First, make sure the host can see the GPU

Run this on your host machine:

nvidia-smi

The useful part of my output was:

NVIDIA-SMI 580.178.04    Driver Version: 580.178.04    CUDA Version: 13.0
GPU 0: NVIDIA GeForce GTX 1060
Memory: 8 MiB / 6144 MiB

If this fails, fix the host driver first. A container cannot rescue a driver that cannot talk to the card.

Give Docker the NVIDIA bridge

I assume Docker is already installed on your host machine. I use sudo docker throughout, so there is no group-membership change or new login to remember.

Next, add NVIDIA's repository and install the Container Toolkit:

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey \
  | sudo gpg --dearmor \
      -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

curl -fsSL https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \
  | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' \
  | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
nvidia-ctk --version

These follow NVIDIA's installation guide. The configuration command updates /etc/docker/daemon.json; restarting Docker makes the daemon read it. On an existing Docker host, schedule that restart around running workloads. Here the daemon was freshly installed.

Let the container introduce itself to the card

Before involving Java, Maven or pricing code, test the bridge:

sudo docker run --rm --gpus all \
  nvidia/cuda:12.9.1-base-ubuntu24.04 nvidia-smi

Inside the container, I saw the same GTX 1060, 580.178.04 driver and 6144 MiB capacity. The driver's CUDA banner still said 13.0. That is expected: nvidia-smi reports the newest CUDA version the host driver supports, not the version installed in the image. The container uses that host driver together with its own CUDA 12.9.1 libraries. A newer driver can run code built with older CUDA versions, so these numbers do not need to match.

--gpus all requests GPU access when the container starts. --rm removes the stopped container; the downloaded image remains available for reuse.

This is an excellent wiring test. It is not yet an application test. nvidia-smi uses the driver's management interface. My Java program also needs to compile and launch a compute kernel. Time to make it earn the green tick.

Build an image that can actually run NablaTensor

Create a directory for the Dockerfile:

mkdir -p ~/gpu-docker-walkthrough
cd ~/gpu-docker-walkthrough

Save this as Dockerfile.gpu. Git, Java and Maven all live in the build container; the Dockerfile fetches the latest source from the default branch:

# syntax=docker/dockerfile:1
FROM maven:3.9.11-eclipse-temurin-25 AS build
RUN git clone --depth=1 https://github.com/nablatensor-dev/nablatensor.git /src
WORKDIR /src
RUN --mount=type=cache,id=nablatensor-maven,target=/root/.m2,sharing=locked \
    mvn -B -q -T 1C -DskipTests install \
 && mkdir -p /out \
 && cp nablatensor-*/target/nablatensor-*.jar /out/

FROM nvidia/cuda:12.9.1-base-ubuntu24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
      cuda-nvrtc-12-9 \
 && rm -rf /var/lib/apt/lists/*
COPY --from=build /opt/java/openjdk /opt/java/openjdk
ENV JAVA_HOME=/opt/java/openjdk
ENV PATH="/opt/java/openjdk/bin:${PATH}"
ENV NVRTC_PATH=/usr/local/cuda-12.9/targets/x86_64-linux/lib/libnvrtc.so.12
COPY --from=build /out/ /app/lib/
WORKDIR /app
ENTRYPOINT ["java", "--enable-native-access=ALL-UNNAMED", "--add-modules", "jdk.incubator.vector", "-cp", "/app/lib/*"]
CMD ["-Dengine=cuda", "-Dscenarios=10000000", "-Dseed=42", "com.nablatensor.examples.BlackScholesBothWays"]

The cache mount requires BuildKit, Docker's modern image builder. Buildx is the command-line plugin I use to drive it. That is a little extra setup, but I don't like watching Maven fetch the same dependencies all over again whenever I rebuild with newer source.

NablaTensor itself has no third-party Java runtime dependencies. Its modules depend on each other, and the only external Java libraries in the project dependencies are test libraries, JUnit and its supporting libraries. Those are used during this build: -DskipTests skips running tests, but Maven still compiles the test sources against them. They are not copied into the final image or used by the pricing program. Maven's test-skipping documentation explains that distinction.

The build machinery has dependencies too. Maven downloads plugins for compiling, processing resources, packaging JARs and installing artifacts, along with the libraries those plugins need. A library with no external runtime dependencies can still give Maven plenty to download.

Leaving out BuildKit and the cache mount would make the Dockerfile look simpler. Unchanged builds could still reuse an entire compiled layer, but refreshing the source would rerun Maven with an empty local repository and fetch those same artifacts again. Keeping the Maven repository in a persistent cache makes repeated builds more efficient.

With Ubuntu's docker.io package, install the Buildx plugin if it is missing:

sudo apt-get install -y docker-buildx
sudo docker buildx version

Then build:

sudo docker buildx build --load -t nablatensor-gpu:walkthrough - < Dockerfile.gpu

This passes just the Dockerfile, with no local filesystem build context. There is no local source checkout to copy or exclude, so no .dockerignore is needed for this recipe.

The first stage fetches the source, compiles the project and collects its JARs. The second keeps the JDK, those JARs and the CUDA runtime compiler; Maven and its download cache stay behind. Tests are skipped during image assembly—I verify the actual workload below.

The Maven cache mount is the useful bit for repeat builds. Docker's BuildKit builder stores /root/.m2 persistently on the host, so downloaded dependencies, plugins and installed Maven artifacts survive a new build. It is managed in Docker's storage, rather than the host user's ~/.m2, and stays out of the final image. Keep using the same builder and sudo invocation to reuse it; pruning build caches can remove it. sharing=locked prevents simultaneous builds from writing to this cache at the same time. Docker's cache mount guide explains how this persists across builds even when a compilation layer changes.

An unchanged Dockerfile can reuse the source checkout and compiled image layer completely. Docker does not check GitHub for new commits when it reuses that layer. To fetch the latest source and compile it again, bypass the build stage's layer cache:

sudo docker buildx build --no-cache-filter build --load \
  -t nablatensor-gpu:walkthrough - < Dockerfile.gpu

The --no-cache-filter option refreshes the named stage. The Maven cache mount survives this refresh, so existing dependencies are reused. -T 1C lets Maven schedule independent modules with one worker per available CPU core, while respecting their dependencies.

I checked both: a second unchanged build reused the compilation layer, and a forced compilation with Maven's -o flag and build-step networking disabled succeeded using the populated cache. No second round of dependency downloads was needed.

NVRTC is the dependency that is easy to miss. NablaTensor generates CUDA C from a recorded valuation and compiles it at runtime. The small CUDA base image does not include that compiler, so I install cuda-nvrtc-12-9 explicitly. NVRTC_PATH identifies the shared library to load. I do not need nvcc or a full CUDA development image here.

Java's --enable-native-access=ALL-UNNAMED enables foreign-function calls into CUDA. The Vector API module is included for the SIMD engine shipped with the examples; the selected CUDA and CPU JIT computations do not use SIMD. Its incubator warning is expected.

No GPU is needed during this build. GPU access enters the story at docker run.

Ten million paths, and some useful answers

Run the image's default command:

sudo docker run --rm --gpus all nablatensor-gpu:walkthrough

The BlackScholesBothWays example prices an option and computes five sensitivities across ten million simulated paths. Give the CPU the same job:

sudo docker run --rm nablatensor-gpu:walkthrough \
  -Dengine=cpu-jit -Dscenarios=10000000 -Dseed=42 \
  com.nablatensor.examples.BlackScholesBothWays

Both returned an option price of about 9.41523. After warm-up, the GPU computed the price and sensitivities in 3.11 ms, versus 386–388 ms on four CPU cores: roughly 124× faster for this workload. Those timings cover the computation, excluding container/JVM startup and compilation. The old gaming card earns its keep.

Prove that CUDA was really selected

A fast result is encouraging. An explicit backend is better evidence. The example accepts -Dengine=cuda, passes it to .on("cuda"), and prints Adjoint Monte-Carlo on cuda. The named engine is required: if unavailable, NablaTensor raises an error instead of silently choosing CPU.

Test that deliberately by omitting GPU access and suppressing the CUDA image's default device enumeration:

sudo docker run --rm -e NVIDIA_VISIBLE_DEVICES=void \
  nablatensor-gpu:walkthrough

On my machine this exited with code 1:

java.lang.IllegalStateException: AAD engine 'cuda' is not usable here;
available: simd, cpu-jit, cpu

That is the failure I want. A CPU fallback would make a broken GPU setup look successful.

For a longer GPU run, watch nvidia-smi in a second terminal on the host:

nvidia-smi -l 1

Very short kernels can finish between samples. An empty utilization snapshot does not outweigh a successfully executed, explicitly selected CUDA backend.

When the GPU plays hide-and-seek

Find which boundary failed before adding more flags:

SymptomWhat to check
Host nvidia-smi failsHost driver and physical device first; Docker comes later
Docker socket permission deniedUse the sudo docker commands above
could not select device driver with GPU capabilitiesToolkit installation, runtime configuration, then daemon restart
Container starts, but CUDA is unavailable--gpus all, compute driver capability, NVRTC installation and NVRTC_PATH
libnvrtc.so cannot be loadedInstall cuda-nvrtc-12-9; check the path against the installed version
Compiler rejects the target architectureToolkit support for the actual GPU; keep this Pascal walkthrough on CUDA 12.9
Image reports an insufficient driverChoose a compatible image/driver pairing; check NVIDIA's release requirements

The CUDA image supplies compute,utility driver capabilities by default. If you override NVIDIA_DRIVER_CAPABILITIES, include both: utility lets nvidia-smi work; compute enables CUDA. A management check can succeed while compute is missing. NVIDIA lists the capabilities in its Docker runtime documentation.

The table is a diagnostic map; it does not claim I encountered every error. The deliberate missing-GPU test above was my failure case. I needed neither --privileged nor hand-written /dev/nvidia* mounts. The toolkit handled the device plumbing.

Pack the application; check the destination

The enjoyable part is how ordinary the final command becomes: docker run --rm --gpus all, followed by the image name. The price, Greeks and convergence checks then look just like they do outside Docker.

The image packages the application and its userspace dependencies. The destination still supplies a compatible driver and GPU. This recipe follows the latest source; the measurements above belong to the revision in the tested-combination table. For exact rebuilds, pin that source revision, base-image digests and package versions instead of following moving targets.

If there's interest, I'll write a follow-up on running the same GPU workload in a Kubernetes cluster.


Questions or corrections? open an issue