Skip to main content

Creating recursion in TPL Dataflow with LinkTo predicates

In the previous post, I showed how to use LinkTo predicates to route messages conditionally across different blocks. Today, we're going to take that concept a step further and do something that surprises most developers the first time they see it:

Link a block back to itself to create recursion — entirely through the dataflow graph, with no explicit recursive method calls.

The core idea

Traditional recursion involves a function calling itself. In TPL Dataflow, we achieve the same result structurally: a block's output is linked back to its own input via a predicate. Messages that match the "recurse" condition loop back, while messages that match the "base case" condition flow forward. The dataflow runtime handles the iteration for us.

Sounds complicated? An example will make it clear immediately.

A good example to illustrate this walks through a directory tree and computing MD5 hashes for every file in a directory. Directories need to be expanded (recursed into), while files need to be processed (hashed).

Here's the dataflow graph we're going to build:


Remark: Notice that the TransformManyBlock is the heart of the recursion. When it receives a folder path, it returns all entries inside that folder. Each entry is then routed by predicate: directories loop back into the same block, and files move forward to be hashed.

Step 1: Define the blocks

Step 2: Wire up the recursive link

This is where the magic happens. We link getFolderContents back to itself with a predicate that matches directories:

The order of these two LinkTo calls matters. TPL Dataflow evaluates predicates in the order the links were created. By placing the directory check first, we ensure subdirectories are always caught before the file check is evaluated.

Step 3: Kick it

We post a single root folder path and call Complete(). The recursion unwinds naturally: as subdirectories are expanded and eventually all entries resolve to files, no new messages loop back, and the block drains on its own.

How the recursion actually terminates

This is the part that confuses people. There is no explicit base-case check or recursion depth limit. Termination relies on two things:

First, the structure of the data guarantees it. A file system is a finite tree. Every directory contains a finite set of entries, and eventually every path resolves to a file (or an empty directory that produces zero output from TransformManyBlock). So the loop-back link naturally stops receiving new messages.

Second, TransformManyBlock returns an empty enumerable for empty directories. When a block produces no output, nothing is posted back, so no new work is generated on that branch. This is the dataflow equivalent of a base case returning immediately.

Why not just use a recursive method?

You might wonder why we'd go through all this instead of a simple recursive method. Here's where the dataflow approach shines:

Parallelism is automatic. Each TransformBlock and TransformManyBlock can process multiple messages concurrently. While one directory is being expanded, other directories at the same level can be expanded simultaneously, and files can already be flowing into the hash computation stage. A naive recursive method processes everything sequentially unless you manually manage threads.

Backpressure is built in. If the hash computation stage can't keep up, the upstream blocks will naturally slow down. You don't need to implement throttling or semaphores yourself.

The pipeline stays flat. No matter how deeply nested the directory tree is, your code doesn't grow a deeper call stack. Each iteration is just another message posted to the block's internal buffer.

Popular posts from this blog

Podman– Command execution failed with exit code 125

After updating WSL on one of the developer machines, Podman failed to work. When we took a look through Podman Desktop, we noticed that Podman had stopped running and returned the following error message: Error: Command execution failed with exit code 125 Here are the steps we tried to fix the issue: We started by running podman info to get some extra details on what could be wrong: >podman info OS: windows/amd64 provider: wsl version: 5.3.1 Cannot connect to Podman. Please verify your connection to the Linux system using `podman system connection list`, or try `podman machine init` and `podman machine start` to manage a new Linux VM Error: unable to connect to Podman socket: failed to connect: dial tcp 127.0.0.1:2655: connectex: No connection could be made because the target machine actively refused it. That makes sense as the podman VM was not running. Let’s check the VM: >podman machine list NAME         ...

Cache stampede: when our cache turned against us

While investigating some performance issues, we ran into an ASP.NET Core API that cached a fairly expensive aggregation query for 60 seconds. Under normal load, that was fine: one request rebuilds the cache, everyone else reads from it. Under peak load, dozens of requests would arrive in that same expiry window, all see a cache miss, and all fire the same expensive query in parallel. The database didn't like that. That was the moment when our caching layer stopped helping and started hurting. A burst of requests comes in at the same time, all miss the cache, and all go hammer the database or the downstream API at once. That's a cache stampede . The cache was supposed to protect our backend, and for a few hundred milliseconds it did the opposite. Why this happens IMemoryCache.GetOrCreate (and its async sibling) looks like it protects you, but it doesn't add any locking on its own. Look at the naive version: public async Task<Report> GetReportAsync(string key) ...

A complex system designed from scratch never works

A few years ago, I worked as an architect on a big mainframe rewrite. I still count it as one of my failures. Not because the technology was wrong, but because I couldn't convince the management team to simplify the approach. Years later, the organization is still struggling to get the new system up and running. I left the project at the time, because I couldn't put my name behind an approach that would take very long and cost a lot of money without a working system to show for it along the way. Gall’s Law That memory keeps coming back to me, because it's a textbook case of Gall's Law playing out in real life. Gall's Law , from John Gall's Systemantics , states it plainly: A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch never works, and it cannot be patched to make it work. You have to start over with a simple system that works. What does that mean in practice,...