# Can I avoid adding duplicate cells to a vtkPolyData?

**URL:** https://discourse.vtk.org/t/can-i-avoid-adding-duplicate-cells-to-a-vtkpolydata/12483
**Category:** Support
**Created:** [October 13, 2023, 1:05pm UTC](https://discourse.vtk.org/t/can-i-avoid-adding-duplicate-cells-to-a-vtkpolydata/12483 "2023-10-13T13:05:47Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![rexthor](https://discourse.vtk.org/user_avatar/discourse.vtk.org/rexthor/32/6481_2.png) [@rexthor](https://discourse.vtk.org/u/rexthor)
#### Post date: [October 13, 2023, 1:05pm UTC](https://discourse.vtk.org/t/can-i-avoid-adding-duplicate-cells-to-a-vtkpolydata/12483/1 "2023-10-13T13:05:48Z")

</div>

I have a vtkUnstructuredGrid composed of vtkHexahedrons (originally it was a vtkPartitionedDataSet composed of vtkStructuredGrids - it made my life easier to just have one partition). The next processing stage needs to store a file that is a list of unique faces present in the volume (I didn’t design the program needing this data - I am just constructing the file).

Originally I was just going to read in the data, append all of the partitions together, create a new vtkPolyData object so that I could run vtkRemoveDuplicatePolys. The problem was that I had 12 million hexahedrons which turn into roughly 72 million quads, and then running the algorithm ate up all of memory and the process gets killed by the OOM reaper.

**One possibility I wanted to explore was using a set of some sort to only add unique faces - I just didn’t know what class to use in the VTK library?**

> **Presently the code (snippet) looks like this:**
>
> ```python
> def polygon_soup(ds):
> "Convert vtkHexahedron (3D cells) in the dataset to faces composed of vtkTriangles."
> logging.debug(f"merging {ds.GetNumberOfPartitions()} partitions in vtkPartitionedDataSet to vtkUnstructuredGrid")
> af = vtk.vtkAppendDataSets()
> af.MergePointsOn()
> for ind in range(ds.GetNumberOfPartitions()):
> af.AddInputData(ds.GetPartition(ind))
> af.Update()
> data = af.GetOutputDataObject(0)
> 
> pd = vtk.vtkPolyData()
> logging.debug(f"saving {data.GetNumberOfPoints()} node locations into new vtkPolyData")
> pd.SetPoints(data.GetPoints())
> 
> logging.debug(f"saving {data.GetPointData().GetNumberOfArrays()} scalar arrays into new vtkPolyData")
> for a_ind in range(data.GetPointData().GetNumberOfArrays()):
> pd.GetPointData().AddArray(data.GetPointData().GetArray(a_ind))
> 
> ca = vtk.vtkCellArray()
> for c_idx in range(ds.GetNumberOfCells()):
> cell = data.GetCell(c_idx)
> for f_idx in range(cell.GetNumberOfFaces()):
> face = cell.GetFace(f_idx)
> node_count = face.GetNumberOfPoints()
> ca.InsertNextCell(node_count)
> for n_idx in range(node_count):
> ca.InsertCellPoint(face.GetPointId(n_idx))
> 
> logging.debug(f"converted {data.GetNumberOfCells()} vtkHexahedrons into {ca.GetNumberOfCells()} vtkQuads")
> pd.SetPolys(ca)
> 
> logging.debug(f"removing duplicate vtkQuads and converting to vtkTriangles")
> cpd = vtk.vtkCleanPolyData()
> cpd.SetInputData(pd)
> rdp = vtk.vtkRemoveDuplicatePolys()
> rdp.SetInputConnection(cpd.GetOutputPort())
> tf = vtk.vtkTriangleFilter()
> tf.SetInputConnection(rdp.GetOutputPort())
> tf.Update()
> return tf.GetOutputDataObject(0)
> 
> ```

---

<div class="post-metadata">

### Author: ![spyridon97](https://discourse.vtk.org/user_avatar/discourse.vtk.org/spyridon97/32/7069_2.png) [@spyridon97](https://discourse.vtk.org/u/spyridon97)
#### Post date: [October 13, 2023, 8:20pm UTC](https://discourse.vtk.org/t/can-i-avoid-adding-duplicate-cells-to-a-vtkpolydata/12483/2 "2023-10-13T20:20:20Z")

</div>

As of now, this is not possible unless you parse all the existing faces already present.

I have implemented a new data structure in VTK named vtkStaticFaceHashLinksTemplate, which vtkGeometryFilter uses to get the external faces. I am envisioning that it could also be used for extracting the internal faces of a mesh, along with internal and external together, which is what you need. I don’t know when I will have some time to work on it, but I hope it will be soon.

---

<div class="post-metadata">

### Author: ![rexthor](https://discourse.vtk.org/user_avatar/discourse.vtk.org/rexthor/32/6481_2.png) [@rexthor](https://discourse.vtk.org/u/rexthor)
#### Post date: [October 18, 2023, 12:15pm UTC](https://discourse.vtk.org/t/can-i-avoid-adding-duplicate-cells-to-a-vtkpolydata/12483/3 "2023-10-18T12:15:12Z")

</div>

Thanks for the reply, I wound up with a crude solution that seems to work (I haven’t done a very detailed check to make sure), but for the benefit of anyone else who stumbles upon this thread, here is how I handled it:

```python
    # ... "data" is a vtkUnstructuredGrid that comes from elided code above.
    NODE_COUNT = data.GetCell(0).GetFace(0).GetNumberOfPoints()
    unique_quads = set()
    for c_idx in range(data.GetNumberOfCells()):
        for f_idx in range(data.GetCell(0).GetNumberOfFaces()):
            face = data.GetCell(c_idx).GetFace(f_idx)
            unique_quads.add(tuple(sorted(face.GetPointId(x) for x in range(NODE_COUNT))))

    ca = vtk.vtkCellArray()
    for q in unique_quads:
        ca.InsertNextCell(NODE_COUNT)
        for n_idx in q:
            ca.InsertCellPoint(n_idx)

    # The cell normals are likely pointing in all sorts of inconvenient
    # directions at this point, but my particular application doesn't care.
    # If your application does care, you will need to fix them.
    data.SetCells(vtk.VTK_QUAD, ca)

```

---

<div class="post-metadata">

### Author: ![dgobbi](https://discourse.vtk.org/user_avatar/discourse.vtk.org/dgobbi/32/18_2.png) [@dgobbi](https://discourse.vtk.org/u/dgobbi)
#### Post date: [October 18, 2023, 7:36pm UTC](https://discourse.vtk.org/t/can-i-avoid-adding-duplicate-cells-to-a-vtkpolydata/12483/4 "2023-10-18T19:36:34Z")

</div>

One thing to watch out for is that sorting the IDs might do more than just flip the normal. It might change the quad from a rectangle into a Z-shape or N-shape.

---

<div class="post-metadata">

### Author: ![rexthor](https://discourse.vtk.org/user_avatar/discourse.vtk.org/rexthor/32/6481_2.png) [@rexthor](https://discourse.vtk.org/u/rexthor)
#### Post date: [October 18, 2023, 8:00pm UTC](https://discourse.vtk.org/t/can-i-avoid-adding-duplicate-cells-to-a-vtkpolydata/12483/5 "2023-10-18T20:00:56Z")

</div>

Good point … instead of sorted I could try:

```python
for f_idx in range(data.GetCell(0).GetNumberOfFaces()):
    face = data.GetCell(c_idx).GetFace(f_idx)
    indices = np.array([face.GetPointId(x) for x in range(FACE_NODE_COUNT)])
    unique_quads.add(tuple(np.roll(indices, np.argmin(indices))))

```

but I don’t think that handles the inverted node order (by which, I mean: `[1, 2, 3, 4]` and `[4, 3, 2, 1]`, which I would only want one of those).

---

<div class="post-metadata">

### Author: ![dgobbi](https://discourse.vtk.org/user_avatar/discourse.vtk.org/dgobbi/32/18_2.png) [@dgobbi](https://discourse.vtk.org/u/dgobbi)
#### Post date: [October 18, 2023, 8:26pm UTC](https://discourse.vtk.org/t/can-i-avoid-adding-duplicate-cells-to-a-vtkpolydata/12483/6 "2023-10-18T20:26:34Z")

</div>

You can use the “set” just for checking for duplicates, and put the original face into your dataset. This also collapses the two loops into just one loop.

```python
    ca = vtk.vtkCellArray()
    unique_quads = set()
    for c_idx in range(data.GetNumberOfCells()):
        for f_idx in range(data.GetCell(0).GetNumberOfFaces()):
            face = data.GetCell(c_idx).GetFace(f_idx)
            faceIDs = [face.GetPointId(x) for x in range(NODE_COUNT)]
            sortedIDs = tuple(sorted(faceIDs))
            if sortedIDs not in unique_quads:
                unique_quads.add(sortedIDs)
                ca.InsertNextCell(len(faceIDs), faceIDs)

```

The `ca.InsertNextCell(len(faceIDs), faceIDs)` is a faster method for inserting cells.
