How to Run Java GPU Workloads in Docker
An old GTX 1060, a fresh Docker install, and ten million Monte Carlo paths: from GPU passthrough to Java option pricing with Greeks.
I'll use a real machine: a small desktop with a GeForce GTX 1060, 6 GB of VRAM, and an Intel i5-7500T. No rented datacenter card, no imaginary terminal output. The finish line is ten million Monte Carlo paths, an option price and five first-order Greeks, computed by NablaTensor inside Docker, then checked against the same computation on the CPU. An old gaming GPU still gets to have a productive afternoon.
Watch the GPU say hello from the host and from Docker, then get to work.
Follow live nvidia-smi telemetry and an ASCII scoreboard built from the measured results. A separate 100-million-path
load run makes GPU activity visible; the CPU/CUDA comparison uses ten million
paths on each backend. Commands are scripted for pacing, and all executions
and readings come from real runs. Long idle waits are shortened in playback.
The three pieces you need
| Piece | Lives where | Its job |
|---|---|---|
| NVIDIA driver | Host | Talks to the physical GPU through the host kernel |
| NVIDIA Container Toolkit | Host | Exposes the requested devices and driver libraries inside Docker |
| Application and CUDA libraries | Image | Runs your code; here, Java plus NVRTC, CUDA's runtime compiler |
The image shares the host's kernel. Installing a kernel driver in a Dockerfile would put it on the wrong side of that arrangement. NVIDIA's toolkit supplies the bridge; its Docker configuration documentation explains device selection and which driver libraries get mounted.
Here is the tested combination:
| Component | Version or hardware |
|---|---|
| Host OS | Ubuntu 24.04.5 LTS, x86-64 |
| CPU | Intel Core i5-7500T, four cores |
| GPU | NVIDIA GeForce GTX 1060, 6 GB, Pascal |
| NVIDIA driver | 580.178.04 |
| Docker Engine | 29.1.3, Ubuntu docker.io package |
| NVIDIA Container Toolkit | 1.20.1 |
| Image CUDA release | 12.9.1 |
| NVRTC package | cuda-nvrtc-12-9, 12.9.86-1 |
| Java | Eclipse Temurin 25.0.1+8, from the Maven build image |
| NablaTensor | Revision 26006a665b074d4e231c8dfa405467afa90bc595, CUDA backend, FP32 |
First, make sure the host can see the GPU
Run this on your host machine:
nvidia-smi
The useful part of my output was:
NVIDIA-SMI 580.178.04 Driver Version: 580.178.04 CUDA Version: 13.0
GPU 0: NVIDIA GeForce GTX 1060
Memory: 8 MiB / 6144 MiB
If this fails, fix the host driver first. A container cannot rescue a driver that cannot talk to the card.
Give Docker the NVIDIA bridge
I assume Docker is already installed on your host machine. I use sudo docker throughout, so there is no group-membership change or
new login to remember.
Next, add NVIDIA's repository and install the Container Toolkit:
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey \
| sudo gpg --dearmor \
-o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -fsSL https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \
| sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' \
| sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
nvidia-ctk --version
These follow NVIDIA's installation guide.
The configuration command updates /etc/docker/daemon.json; restarting
Docker makes the daemon read it. On an existing Docker host, schedule that
restart around running workloads. Here the daemon was freshly installed.
Let the container introduce itself to the card
Before involving Java, Maven or pricing code, test the bridge:
sudo docker run --rm --gpus all \
nvidia/cuda:12.9.1-base-ubuntu24.04 nvidia-smi
Inside the container, I saw the same GTX 1060, 580.178.04 driver and
6144 MiB capacity. The driver's CUDA banner still said 13.0. That is
expected: nvidia-smi reports the newest CUDA version the host driver
supports, not the version installed in the image. The container uses
that host driver together with its own CUDA 12.9.1 libraries. A newer
driver can run code built with older CUDA versions, so these numbers do
not need to match.
--gpus all requests GPU access when the container starts. --rm removes
the stopped container; the downloaded image remains available for reuse.
This is an excellent wiring test. It is not yet an application test.
nvidia-smi uses the driver's management interface. My Java program also
needs to compile and launch a compute kernel. Time to make it earn the
green tick.
Build an image that can actually run NablaTensor
Create a directory for the Dockerfile:
mkdir -p ~/gpu-docker-walkthrough
cd ~/gpu-docker-walkthrough
Save this as Dockerfile.gpu. Git, Java and Maven all live in the build
container; the Dockerfile fetches the latest source from the default branch:
# syntax=docker/dockerfile:1
FROM maven:3.9.11-eclipse-temurin-25 AS build
RUN git clone --depth=1 https://github.com/nablatensor-dev/nablatensor.git /src
WORKDIR /src
RUN --mount=type=cache,id=nablatensor-maven,target=/root/.m2,sharing=locked \
mvn -B -q -T 1C -DskipTests install \
&& mkdir -p /out \
&& cp nablatensor-*/target/nablatensor-*.jar /out/
FROM nvidia/cuda:12.9.1-base-ubuntu24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
cuda-nvrtc-12-9 \
&& rm -rf /var/lib/apt/lists/*
COPY --from=build /opt/java/openjdk /opt/java/openjdk
ENV JAVA_HOME=/opt/java/openjdk
ENV PATH="/opt/java/openjdk/bin:${PATH}"
ENV NVRTC_PATH=/usr/local/cuda-12.9/targets/x86_64-linux/lib/libnvrtc.so.12
COPY --from=build /out/ /app/lib/
WORKDIR /app
ENTRYPOINT ["java", "--enable-native-access=ALL-UNNAMED", "--add-modules", "jdk.incubator.vector", "-cp", "/app/lib/*"]
CMD ["-Dengine=cuda", "-Dscenarios=10000000", "-Dseed=42", "com.nablatensor.examples.BlackScholesBothWays"]
The cache mount requires BuildKit, Docker's modern image builder. Buildx is the command-line plugin I use to drive it. That is a little extra setup, but I don't like watching Maven fetch the same dependencies all over again whenever I rebuild with newer source.
NablaTensor itself has no third-party Java runtime dependencies. Its
modules depend on each other, and the only external Java libraries in
the project dependencies are test libraries, JUnit and its supporting
libraries. Those are used during this build: -DskipTests skips running
tests, but Maven still compiles the test sources against them. They are
not copied into the final image or used by the pricing program. Maven's
test-skipping documentation
explains that distinction.
The build machinery has dependencies too. Maven downloads plugins for compiling, processing resources, packaging JARs and installing artifacts, along with the libraries those plugins need. A library with no external runtime dependencies can still give Maven plenty to download.
Leaving out BuildKit and the cache mount would make the Dockerfile look simpler. Unchanged builds could still reuse an entire compiled layer, but refreshing the source would rerun Maven with an empty local repository and fetch those same artifacts again. Keeping the Maven repository in a persistent cache makes repeated builds more efficient.
With Ubuntu's docker.io package, install the Buildx plugin if it is missing:
sudo apt-get install -y docker-buildx
sudo docker buildx version
Then build:
sudo docker buildx build --load -t nablatensor-gpu:walkthrough - < Dockerfile.gpu
This passes just the Dockerfile, with no local filesystem build context.
There is no local source checkout to copy or exclude, so no .dockerignore
is needed for this recipe.
The first stage fetches the source, compiles the project and collects its JARs. The second keeps the JDK, those JARs and the CUDA runtime compiler; Maven and its download cache stay behind. Tests are skipped during image assembly—I verify the actual workload below.
The Maven cache mount is the useful bit for repeat builds. Docker's
BuildKit builder stores /root/.m2 persistently on the host, so downloaded
dependencies, plugins and installed Maven artifacts survive a new build.
It is managed in Docker's storage, rather than the host user's
~/.m2, and stays out of the final image. Keep using the same builder and
sudo invocation to reuse it; pruning build caches can remove it.
sharing=locked prevents simultaneous builds from writing to this cache
at the same time.
Docker's cache mount guide
explains how this persists across builds even when a compilation layer changes.
An unchanged Dockerfile can reuse the source checkout and compiled image layer completely. Docker does not check GitHub for new commits when it reuses that layer. To fetch the latest source and compile it again, bypass the build stage's layer cache:
sudo docker buildx build --no-cache-filter build --load \
-t nablatensor-gpu:walkthrough - < Dockerfile.gpu
The --no-cache-filter option
refreshes the named stage. The Maven cache mount survives this refresh, so existing dependencies
are reused. -T 1C lets Maven schedule independent modules with one
worker per available CPU core, while respecting their dependencies.
I checked both: a second unchanged build reused the compilation
layer, and a forced compilation with Maven's -o flag and build-step
networking disabled succeeded using the populated cache. No second round
of dependency downloads was needed.
NVRTC is the dependency that is easy to miss. NablaTensor generates
CUDA C from a recorded valuation and compiles it at runtime. The small
CUDA base image does not include that compiler, so I install
cuda-nvrtc-12-9 explicitly. NVRTC_PATH identifies the shared library to
load. I do not need nvcc or a full CUDA development image here.
Java's --enable-native-access=ALL-UNNAMED enables foreign-function calls
into CUDA. The Vector API module is included for the SIMD engine shipped
with the examples; the selected CUDA and CPU JIT computations do not use
SIMD. Its incubator warning is expected.
No GPU is needed during this build. GPU access enters the story at
docker run.
Ten million paths, and some useful answers
Run the image's default command:
sudo docker run --rm --gpus all nablatensor-gpu:walkthrough
The BlackScholesBothWays example
prices an option and computes five sensitivities across ten million simulated paths.
Give the CPU the same job:
sudo docker run --rm nablatensor-gpu:walkthrough \
-Dengine=cpu-jit -Dscenarios=10000000 -Dseed=42 \
com.nablatensor.examples.BlackScholesBothWays
Both returned an option price of about 9.41523. After warm-up, the GPU computed the price and sensitivities in 3.11 ms, versus 386–388 ms on four CPU cores: roughly 124× faster for this workload. Those timings cover the computation, excluding container/JVM startup and compilation. The old gaming card earns its keep.
Prove that CUDA was really selected
A fast result is encouraging. An explicit backend is better evidence.
The example accepts -Dengine=cuda, passes it to .on("cuda"), and prints
Adjoint Monte-Carlo on cuda. The named engine is required: if unavailable,
NablaTensor raises an error instead of silently choosing CPU.
Test that deliberately by omitting GPU access and suppressing the CUDA image's default device enumeration:
sudo docker run --rm -e NVIDIA_VISIBLE_DEVICES=void \
nablatensor-gpu:walkthrough
On my machine this exited with code 1:
java.lang.IllegalStateException: AAD engine 'cuda' is not usable here;
available: simd, cpu-jit, cpu
That is the failure I want. A CPU fallback would make a broken GPU setup look successful.
For a longer GPU run, watch nvidia-smi in a second terminal on the host:
nvidia-smi -l 1
Very short kernels can finish between samples. An empty utilization snapshot does not outweigh a successfully executed, explicitly selected CUDA backend.
When the GPU plays hide-and-seek
Find which boundary failed before adding more flags:
| Symptom | What to check |
|---|---|
Host nvidia-smi fails | Host driver and physical device first; Docker comes later |
| Docker socket permission denied | Use the sudo docker commands above |
could not select device driver with GPU capabilities | Toolkit installation, runtime configuration, then daemon restart |
| Container starts, but CUDA is unavailable | --gpus all, compute driver capability, NVRTC installation and NVRTC_PATH |
libnvrtc.so cannot be loaded | Install cuda-nvrtc-12-9; check the path against the installed version |
| Compiler rejects the target architecture | Toolkit support for the actual GPU; keep this Pascal walkthrough on CUDA 12.9 |
| Image reports an insufficient driver | Choose a compatible image/driver pairing; check NVIDIA's release requirements |
The CUDA image supplies compute,utility driver capabilities by default.
If you override NVIDIA_DRIVER_CAPABILITIES, include both: utility lets
nvidia-smi work; compute enables CUDA. A management check can succeed
while compute is missing. NVIDIA lists the capabilities in its
Docker runtime documentation.
The table is a diagnostic map; it does not claim I encountered every
error. The deliberate missing-GPU test above was my failure case.
I needed neither --privileged nor hand-written /dev/nvidia* mounts.
The toolkit handled the device plumbing.
Pack the application; check the destination
The enjoyable part is how ordinary the final command becomes:
docker run --rm --gpus all, followed by the image name. The price,
Greeks and convergence checks then look just like they do outside Docker.
The image packages the application and its userspace dependencies. The destination still supplies a compatible driver and GPU. This recipe follows the latest source; the measurements above belong to the revision in the tested-combination table. For exact rebuilds, pin that source revision, base-image digests and package versions instead of following moving targets.
If there's interest, I'll write a follow-up on running the same GPU workload in a Kubernetes cluster.
