# Using \`numpy\_to\_vtk\` I can't seem to get arrays to stay in dataset?

**URL:** https://discourse.vtk.org/t/using-numpy-to-vtk-i-cant-seem-to-get-arrays-to-stay-in-dataset/12425
**Category:** Support
**Created:** [October 5, 2023, 2:56pm UTC](https://discourse.vtk.org/t/using-numpy-to-vtk-i-cant-seem-to-get-arrays-to-stay-in-dataset/12425 "2023-10-05T14:56:51Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![rexthor](https://discourse.vtk.org/user_avatar/discourse.vtk.org/rexthor/32/6481_2.png) [@rexthor](https://discourse.vtk.org/u/rexthor)
#### Post date: [October 5, 2023, 2:56pm UTC](https://discourse.vtk.org/t/using-numpy-to-vtk-i-cant-seem-to-get-arrays-to-stay-in-dataset/12425/1 "2023-10-05T14:56:51Z")

</div>

I think that there are issues I am having with data starting in NumPy, but not properly “sticking” in my dataset. Peripherally discussed [here](https://discourse.vtk.org/t/appending-data-field-to-a-vtk-file-python-vtk/3220/4). The data I am trying to put into VTK lives in a few very large files. One has geometry, and another has scalar values associated with each node (temperature, pressure, etc.)

> **Schematically, I start out with a \`vtkPartitionedDataSet\`, and \*in a function\* I add each block of the geometry to it.**
>
> ```python
> def get_geometry(fname, pdset):
> with open(fname, "rb") as f:
> # find blocks
> for idx, b in enumerate(blocks):
> sg = vtk.vtkStructuredGrid()
> # ...
> pdset.SetPartition(idx, sg)
> 
> ```
> 
> The geometry creation looks like it works well - I can save the file and plot it in ParaView, it looks fine.

> **Then, I read in the scalar data as an array of \`vtkFloatArray\`:**
>
> ```python
> def get_scalars(fname):
> "Return list-of-list-of variables."
> solution_data = []
> with open(fname, "rb") as f:
> # ... find blocks
> for idx, b in enumerate(blocks):
> for _ in range(number_of_variables):
> d = np.fromfile(f, dtype=">f4", count=elementcount)
> d = d.reshape(size_of_block)
> solution_data.append(d.ravel().byteswap())
> 
> return solution_data
> 
> ```

> **And in yet another function, I (try to) add the data to the dataset:**
>
> ```python
> def store_scalars(dataset):
> variables = get_scalars("huge_file.dat")
> namelist = get_namelist("name_list.txt")
> 
> for j, g in enumerate(variables):
> p = data.GetPartition(j)
> for k, v in enumerate(g):
> a = numpy_to_vtk(v)
> a.SetName(namelist[k])
> 
> logging.debug(f"vtkFloatArray.GetValueRange(): {a.GetValueRange()}")
> p.GetPointData().AddArray(a)
> p.Modified()
> 
> ```

So adding the data, the debug statement that prints out the variable array range is nominally correct. Max and min values are about what I would expect.

But in a debug print just before I save the file to disk, the values are not correct? In short, I create `vtkFloatArray`s to store in the dataset, and a check at creation time shows that the arrays seem to have the right information in them. But, seconds later, when I make a check of the array contents before saving them to disk, they seem to contain garbage … (like numbers on the order of 1e38, and NaN … perhaps some random chunk of memory after a garbage collection? And I get those numbers instead of something worse like a segfault?)

Is there some way to deep copy the data out of the NumPy allocated memory and store them more permenantly inside of the `vtkStructuredGrid` partitions?

---

<div class="post-metadata">

### Author: ![rexthor](https://discourse.vtk.org/user_avatar/discourse.vtk.org/rexthor/32/6481_2.png) [@rexthor](https://discourse.vtk.org/u/rexthor)
#### Post date: [October 5, 2023, 6:06pm UTC](https://discourse.vtk.org/t/using-numpy-to-vtk-i-cant-seem-to-get-arrays-to-stay-in-dataset/12425/2 "2023-10-05T18:06:32Z")

</div>

It seemed that the obvious answer was to not have NumPy allocate the data memory.  
This is related to [a problem I had earlier](https://discourse.vtk.org/t/problems-using-vtkfloatarray/12416).

So, rather than:

```python
a = numpy_to_vtk(large_NumPy_array)
dataset.GetPointData().AddArray(a)
dataset.Modified()

```

I did this:

```python
a = vtk.vtkFloatArray()
a.SetNumberOfComponents(1)
for e in large_NumPy_array:
    a.InsertInextValue(e)

dataset.GetPointData().AddArray(a)
dataset.Modified()

```
