GPU volume rendering reads the wrong memory for partitioned volumes larger than 2³¹ voxels (int overflow in vtkVolumeTexture)

Hello, I’m a 3D Slicer user and ran into this VTK issue with volume rendering on large-ish volumes. Although I was a SW dev in a past life, I used Claude Opus 5.5 to debug this, and the post below is basically from Claude. The test case here is straight VTK and doesn’t rely on Slicer. Please let me know what you think - Thank you!

Summary

When vtkOpenGLGPUVolumeRayCastMapper splits a volume into blocks (SetPartitions), vtkVolumeTexture::CreateBlocks computes each block’s starting tuple index in 32-bit int. For volumes with more than 2³¹ voxels, the index of later blocks overflows and wraps negative. VTK then uploads those blocks from memory up to ~2 GB before the scalar array. Depending on what is mapped at that address, this either crashes inside the OpenGL driver or silently renders garbage.

This is still present on master (checked at b7eead4d07) and in the 9.7.1 wheel from PyPI.

The code

Rendering/VolumeOpenGL2/vtkVolumeTexture.cxx, in CreateBlocks:

block->TupleIndex =
  ext[4] * this->FullSize[0] * this->FullSize[1] + ext[2] * this->FullSize[0] + ext[0];

ext and FullSize are vtkTuple<int, N>, so the whole expression is evaluated in int and only then widened to vtkIdType. On the 9.7.1 macOS arm64 wheel, the compiled code is a 32-bit madd followed by sxtw, i.e. the wrapped value is sign-extended.

How I hit it

3D Slicer partitions any volume axis longer than 2048 on macOS. Volume rendering a 1671 × 1473 × 2813 unsigned short micro-CT scan crashed with EXC_BAD_ACCESS in vtkTextureObject::Create3DFromRaw → glTexImage3D → Apple’s Metal driver. The second z-block starts at z = 1406, so its correct index is 1406 × 1671 × 1473 = 3,460,704,498. That wraps to −834,262,798, and the faulting address in the crash was exactly the array base minus 2 × 834,262,798 bytes.

The bug isn’t platform-specific, but it only affects partitioned volumes. Slicer partitions at 2048 voxels per axis on macOS and 4096 elsewhere, so in practice macOS users hit it first. The same volume on Windows (NVIDIA Quadro RTX 5000) uploads as a single texture and renders fine. Forcing partitions with SetPartitions reproduces it on any platform.

Reproducer (pure VTK Python)

A 2048 × 2048 × 600 unsigned char volume (~2.5 GB RAM) contains one small bright cube. It is rendered with MIP twice, as a single texture and split into 8 z-blocks, and the lit pixels are counted. The two renders should give the same count. The last block starts at z0 = 524, and 524 × 2048 × 2048 = 2,197,815,296 > 2³¹ − 1.

import numpy as np
from vtkmodules.util.numpy_support import numpy_to_vtk, vtk_to_numpy
import vtkmodules.vtkRenderingOpenGL2  # noqa: F401
from vtkmodules.vtkCommonDataModel import vtkImageData, vtkPiecewiseFunction
from vtkmodules.vtkRenderingCore import (
    vtkColorTransferFunction, vtkRenderer, vtkRenderWindow,
    vtkVolume, vtkVolumeProperty, vtkWindowToImageFilter,
)
from vtkmodules.vtkRenderingVolumeOpenGL2 import vtkOpenGLGPUVolumeRayCastMapper

nx, ny, nz = 2048, 2048, 600
scalars = np.zeros((nz, ny, nx), dtype=np.uint8)
scalars[10:60, 900:1100, 900:1100] = 255

image = vtkImageData()
image.SetDimensions(nx, ny, nz)
image.GetPointData().SetScalars(numpy_to_vtk(scalars.ravel(), deep=False))

opacity = vtkPiecewiseFunction()
opacity.AddPoint(0, 0.0)
opacity.AddPoint(255, 1.0)
color = vtkColorTransferFunction()
color.AddRGBPoint(0, 0.0, 0.0, 0.0)
color.AddRGBPoint(255, 1.0, 1.0, 1.0)
prop = vtkVolumeProperty()
prop.SetScalarOpacity(opacity)
prop.SetColor(color)


def count_lit_pixels(partitions):
    mapper = vtkOpenGLGPUVolumeRayCastMapper()
    mapper.SetInputData(image)
    mapper.SetBlendModeToMaximumIntensity()
    mapper.SetPartitions(*partitions)
    volume = vtkVolume()
    volume.SetMapper(mapper)
    volume.SetProperty(prop)
    renderer = vtkRenderer()
    renderer.AddVolume(volume)
    window = vtkRenderWindow()
    window.AddRenderer(renderer)
    window.SetSize(400, 400)
    renderer.ResetCamera()
    window.Render()
    grab = vtkWindowToImageFilter()
    grab.SetInput(window)
    grab.Update()
    pixels = vtk_to_numpy(grab.GetOutput().GetPointData().GetScalars())
    window.Finalize()
    return np.count_nonzero(pixels.max(axis=1))


print("lit pixels, single block :", count_lit_pixels((1, 1, 1)))
print("lit pixels, 8 z-blocks   :", count_lit_pixels((1, 1, 8)))

Results on an Apple M5 Max, macOS 26.5:

single block 8 z-blocks
VTK 9.7.1 (PyPI wheel) 676 79,524
VTK 9.6.1 + fix below 676 676

Under lldb on the 9.7.1 wheel, blocks 1–7 upload from the expected addresses, and block 8 uploads from exactly 2 GiB below the array base. Here that address was in the dyld shared cache, which is readable, so the result was garbage rather than a crash. On other systems or with other sizes it may segfault instead.

Proposed fix

Do the arithmetic in vtkIdType. The kInc product in the HandleLargeDataTypes path of LoadTexture has the same pattern. It only overflows for a single slice of more than 2³¹ voxels, so it’s very unlikely in practice, but it seems worth fixing at the same time:

--- a/Rendering/VolumeOpenGL2/vtkVolumeTexture.cxx
+++ b/Rendering/VolumeOpenGL2/vtkVolumeTexture.cxx
@@ -246,8 +246,10 @@ void vtkVolumeTexture::CreateBlocks(unsigned int format, unsigned int internalFo
 
     // Compute tuple index (array aligned in x -> Y -> Z)
     // index = z0 * Dx * Dy + y0 * Dx + x0
-    block->TupleIndex =
-      ext[4] * this->FullSize[0] * this->FullSize[1] + ext[2] * this->FullSize[0] + ext[0];
+    // Use vtkIdType arithmetic: the product overflows int for volumes with
+    // more than 2^31 voxels.
+    block->TupleIndex = static_cast<vtkIdType>(ext[4]) * this->FullSize[0] * this->FullSize[1] +
+      static_cast<vtkIdType>(ext[2]) * this->FullSize[0] + ext[0];
 
@@ -395,7 +397,7 @@ bool vtkVolumeTexture::LoadTexture(int interpolation, VolumeBlock* volBlock)
-    vtkIdType const kInc = this->FullSize[0] * this->FullSize[1];
+    vtkIdType const kInc = static_cast<vtkIdType>(this->FullSize[0]) * this->FullSize[1];

The patch applies cleanly to current master. I’ve been running it in a local Slicer build, where it fixes both the original crash and the reproducer above.

If this looks right, I’m happy to open a merge request. Would you like a regression test along the lines of the reproducer, or is ~2.5 GB too much memory for the test suite?