# About proper distributed AbortCheck

**URL:** https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315
**Category:** Development
**Created:** [March 5, 2026, 3:56pm UTC](https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315 "2026-03-05T15:56:35Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![mwestphal](https://discourse.vtk.org/user_avatar/discourse.vtk.org/mwestphal/32/19_2.png) [@mwestphal](https://discourse.vtk.org/u/mwestphal)
#### Post date: [March 5, 2026, 3:56pm UTC](https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315/1 "2026-03-05T15:56:35Z")

</div>

1. **Context**

A few years ago, @Stephen_Crowell @Christos_Tsolakis and @berk.geveci led an effort to improve abort mechanism in VTK, this effort was not fully integrated but great point and design were made here: [https://gitlab.kitware.com/vtk/vtk/-/issues/18463](https://gitlab.kitware.com/vtk/vtk/-/issues/18463) (must read).

Im picking this up again, specifically I’m looking into the possibility to abort a filter while it is running in a distributed pipeline.

**2. The `CheckAbort` problem**

With the current implementation of `CheckAbort()`, the idea is to magically position the `AbortExecute` flag on all distributed processes **at the same time**. Then whenever each process use `CheckAbort()` on their own time, they will abort and return in error.

Since distributed processing in VTK is implemented using MPI, this is just not possible, as there is no way to talk to a distributed process while it is engaged in communication with another process sending messages back and forth.

Moreover, even in the scenario were we could position a flag on all processes at the same time, then there are still cases where it would not work, eg:

tick 0: proc A is processing, proc B is processing  
tick 1: proc A is waiting on a MPI receive, proc B is doing some heavy processing  
tick 2: all abort flags are set on all processes, proc A is still waiting on a MPI receive, proc B is still doing some heavy processing  
tick 3:proc A is still waiting on a MPI receive, proc B finished processing, and check abort flag  
tick 4: proc A is still waiting on a MPI receive, proc B aborts everything and never send with MPI  
tick 5: proc A is still waiting on a MPI receive for a message that will never come, **dealock**

So clearly, current `CheckAbort()` mechanism doesn’t fit distributed computing, we need a way to synchronize these checks.

**3. Two main usecases**

There is two usecases to account for, MPI using filters, and non-MPI using filters.

For MPI-using filters, there is no magic possibility, the CheckAbort calls **MUST** be synchronous and akin to a MPIBarrier call. All processes wait for each other and then reduce the abort state so that everyone abort at the same time if any processes was aborted.

For non-MPI using filters, there can be more leniency, as the CheckAbort and reduction of the abort state can be started at different point on different processes, they will ultimely wait for each other and then abort together.

But there is one problem. processes don’t wait around, a singular processes with no data to process (because of a clip at the beginning of the pipeline for example) would just keep going and finish updating fully while the other processes are still clipping!

There is no other choice then to synchronize at the end of processing of each filter, before moving to the next filter.

It is not ideal in a task based distributed computing system but in reality, with VTK pipeline, processes with not a lot of work with one filter tends to not have a lot of work with all filters, so there is no balancing to do.

**4. Actual Implementation**

So we need two implementation, a synchronous impl and a lazy impl, and these will be triggered by calls from within the filter implementation. The lazy impl also MUST not depend on anything but Common modules in VTK because we cannot add dependency to ParallelCore.

So this must be handled using events.

`vtkAlgorithm::CheckAbortAndInvoke` that invokes `CHECK_ABORT` event and then actually `CheckAbort` is pretty easy, as long as it documents the needs for synchronous calls on MPI filters.

Adding `vtkAlgorithm::CheckAbortDone()` that invoke a new `CHECK_ABORT_DONE` event will then takes care of the synchronization of the end of filters.

**5. What about the inter process communication ?**

This is where it is still a bit murky to me.  
I could easilly observe these events from the ParaView server layer and react on it to do the actual abort state reduction / syncronisation but I feel like this should be provided by VTK somehow, and I don’t know exactly how this could look like.

A possibility could be an AbortObserver that would be added to vtkAlgorithm in a generic version and ParallelMPI would factory-provide a specialized version that would be responsible to reducing the abort. This is more or less how the vtkProgressObserver is implemented.

Let me know what you think!

---

<div class="post-metadata">

### Author: ![dcthomp](https://discourse.vtk.org/user_avatar/discourse.vtk.org/dcthomp/32/60_2.png) [@dcthomp](https://discourse.vtk.org/u/dcthomp)
#### Post date: [March 6, 2026, 5:24pm UTC](https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315/2 "2026-03-06T17:24:32Z")

</div>

@mwestphal Have you thought about filters that run other filters internally? Could that cause deadlocks in the same way? Especially when processing distributed datasets that may not have cells on every rank?

---

<div class="post-metadata">

### Author: ![mwestphal](https://discourse.vtk.org/user_avatar/discourse.vtk.org/mwestphal/32/19_2.png) [@mwestphal](https://discourse.vtk.org/u/mwestphal)
#### Post date: [March 9, 2026, 8:22am UTC](https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315/3 "2026-03-09T08:22:52Z")

</div>

Hi @dcthomp

Good question!

I though about it and it seems fine, let me develop.

There are two ways to integrate a filter inside another filter, either through SetInputData, or SetInputConnection.

SetInputData is straightforward, the CheckAbortEvent system will not be listening to this internal filter, which means that it will behave as if its a simple CheckAbort, which will just not abort at all because the flag is not set.  
Of course, the high level filter will need to perform proper CheckAbort to be abortable, as explained in my message above.  
Actually forwarding the CheckAbortEvent and aborting lower level filter would actually be possible, but it would some manual coce and the proposition above doesnt provide tools for that specifically.

SetInputConnection is mot complex because the SetAbortExecuteAndUpdateTime not only put the abort flag on the filter but also on downstream filter as explained in Stepen issue linked above.  
This means the internal filter will abort as well. In distributed context, it may means that only the rank 0 will abort unless the abort flag reduction takes place. For the abort flag reduction to take place, then it means the filter will need to implement proper AbortCheck forwarding out AND in.  
But if not implemented, it will only means the other ranks won’t be aborting right away, which is not a big deal UNLESS this is a MPI filter that expect communication to happen.

In any case, this is a very niche usecase that is also currently not supported any better, with the classic CheckAbort, we would have the same behavior, just no tooling at all to reduce the abort flag. With my proposition, there is a litle bit of tooling for filter developpers to use to try to make this usecase work.

I hope that answers your question!

---

<div class="post-metadata">

### Author: ![dcthomp](https://discourse.vtk.org/user_avatar/discourse.vtk.org/dcthomp/32/60_2.png) [@dcthomp](https://discourse.vtk.org/u/dcthomp)
#### Post date: [March 21, 2026, 2:44pm UTC](https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315/4 "2026-03-21T14:44:08Z")

</div>

@mwestphal Thanks, I think that covers most of my concerns. If there was a way to detect that an internal filter was configured with `SetInputConnection()` to the “external” filter’s input, that would be nice (because we could at least warn that MPI communication could cause deadlocks during aborts in that case). But I don’t think there is a way to do that. But if the documentation doesn’t already say not to use `SetInputConnection()` inside `RequestData()` implementations, it should.

---

<div class="post-metadata">

### Author: ![mwestphal](https://discourse.vtk.org/user_avatar/discourse.vtk.org/mwestphal/32/19_2.png) [@mwestphal](https://discourse.vtk.org/u/mwestphal)
#### Post date: [March 23, 2026, 8:35am UTC](https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315/5 "2026-03-23T08:35:56Z")

</div>

> [@dcthomp](#):
>
> But if the documentation doesn’t already say not to use `SetInputConnection()` inside `RequestData()` implementations, it should.

It should not, this is a valid workflow, albeit rarely used.

---

<div class="post-metadata">

### Author: ![nicolas.vuaille](https://discourse.vtk.org/user_avatar/discourse.vtk.org/nicolas.vuaille/32/5363_2.png) [@nicolas.vuaille](https://discourse.vtk.org/u/nicolas.vuaille)
#### Post date: [March 24, 2026, 6:55am UTC](https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315/6 "2026-03-24T06:55:23Z")

</div>

I would like to mitigate:

> [@mwestphal](#):
>
> this is a valid workflow,

I do not see a case where it is the only, good, way to go. But I can see unwanted side effects. So IMO we should discourage this pattern. But it may be a discussion for another place.

---

<div class="post-metadata">

### Author: ![mwestphal](https://discourse.vtk.org/user_avatar/discourse.vtk.org/mwestphal/32/19_2.png) [@mwestphal](https://discourse.vtk.org/u/mwestphal)
#### Post date: [August 26, 2026, 8:38am UTC](https://discourse.vtk.org/t/about-proper-distributed-abortcheck/16315/7 "2026-08-26T08:38:25Z")

</div>

Duplicating my answer here: [Improving Abort Functionality in VTK](https://discourse.vtk.org/t/improving-abort-functionality-in-vtk/7813)

I’m afraid it may not be a nice as he would have hoped but here it is (read [https://gitlab.kitware.com/vtk/vtk/-/work\_items/18463](https://gitlab.kitware.com/vtk/vtk/-/work_items/18463) first) :

The current abort implementation is merely a `CheckAbort` being repeatedly called at high velocity in the filters. This is problematic because there is no way to call `SetAbortExecute` while the filter execute in a distributed way.

Applications however rely on the rank 0 sending progress events to track progress of filters, and executing code on rank 0 during a progress event is possible, which of course include calling `SetAbortExecute`.

So once rank 0 has the abort flag, how to transmit it to other ranks ?

Well, the idea is of course to send it from rank 0 to other ranks, **but it is not VTK responsability to do that.**

VTK is merely responsible to signal to applications that each rank is currently checking the abort flag.

So I will introduce a `CheckAbortAndInvoke` method, that will invoke a `CheckAbort` event and then check if the abort is set.

This is the first hurdle, but then, since application uses this event to trigger multiprocess communication, we end up needing to call `CheckAbortAndInvoke` the exact same number of time on each rank, which is an unreassonnable ask.

The solution is to use another signal, at the end of the request data implementation, which can be emitted easilly using `CleanupAbortCheckEvent`.

Application are now able to cleanup any multi process communication setup that they have been using during the CheckAbort event handling.

I was able to implement that in ParaView using MPI NoBlockSend/NoBlockReceive/Test/Cancel API : [https://gitlab.kitware.com/paraview/paraview/-/merge\_requests/7921](https://gitlab.kitware.com/paraview/paraview/-/merge_requests/7921)

Of course, it means that any VTK user willing to get the same abort support also require to implement the communication, but I’m afraid this implementation is outside of the scope of vtkAlgorithm responsabilities.
