Skip to main content

.NET Aspire: The price of forgetting WithReference

Recently I lost way more time than I'd like to admit on an error that turned out to be one missing line of code. The symptom looked like a networking problem. The cause was a missing WithReference call in my Aspire AppHost.

Here's the exception my proxy threw the moment it tried to forward a request to my API:

System.Net.Http.HttpRequestException: No such host is known. (api:443)
 ---> System.Net.Sockets.SocketException (11001): No such host is known.
   at System.Net.Sockets.Socket.AwaitableSocketAsyncEventArgs.ThrowException(SocketError error, CancellationToken cancellationToken)
   at System.Net.Sockets.Socket.AwaitableSocketAsyncEventArgs.System.Threading.Tasks.Sources.IValueTaskSource.GetResult(Int16 token)
   at System.Net.Http.HttpConnectionPool.ConnectToTcpHostAsync(String host, Int32 port, HttpRequestMessage initialRequest, Boolean async, CancellationToken cancellationToken)
   --- End of inner exception stack trace ---
   at System.Net.Http.HttpConnectionPool.ConnectToTcpHostAsync(String host, Int32 port, HttpRequestMessage initialRequest, Boolean async, CancellationToken cancellationToken)
   at System.Net.Http.HttpConnectionPool.ConnectAsync(HttpRequestMessage request, Boolean async, CancellationToken cancellationToken)
   at System.Net.Http.HttpConnectionPool.InjectNewHttp2ConnectionAsync(QueueItem queueItem)
   at System.Threading.Tasks.TaskCompletionSourceWithCancellation`1.WaitWithCancellationAsync(CancellationToken cancellationToken)
   at System.Net.Http.HttpConnectionWaiter`1.WaitForConnectionWithTelemetryAsync(HttpRequestMessage request, HttpConnectionPool pool, Boolean async, CancellationToken requestCancellationToken)
   at System.Net.Http.HttpConnectionPool.SendWithVersionDetectionAndRetryAsync(HttpRequestMessage request, Boolean async, Boolean doRequestAuth, CancellationToken cancellationToken)
   at System.Net.Http.Metrics.MetricsHandler.SendAsyncWithMetrics(HttpRequestMessage request, Boolean async, CancellationToken cancellationToken)
   at System.Net.Http.DiagnosticsHandler.SendAsyncCore(HttpRequestMessage request, Boolean async, CancellationToken cancellationToken)
   at System.Net.Http.SocketsHttpHandler.<SendAsync>g__CreateHandlerAndSendAsync|115_0(HttpRequestMessage request, CancellationToken cancellationToken)
   at Yarp.ReverseProxy.Forwarder.HttpForwarder.SendAsync(HttpContext context, String destinationPrefix, HttpMessageInvoker httpClient, ForwarderRequestConfig requestConfig, HttpTransformer transformer, CancellationToken cancellationToken)

A DNS lookup failing on the literal string api. Not api.something.internal, just api. That should have been the giveaway right away, but when you're staring at a YARP stack trace at the end of the day, it isn't.

The setup

I have a YARP-based reverse proxy in front of an API, both wired up through Aspire:

var api = builder.AddProject<Projects.Demo_Api>("api")
    .WithExternalHttpEndpoints();

var proxy = builder.AddProject<Projects.Demo_Proxy>("proxy")
    .WaitFor(api)
    .WithExternalHttpEndpoints();

Looks reasonable. The proxy's YARP config points at a destination named api, Aspire starts both projects fine, the dashboard shows both as healthy. And then the first request through the proxy blows up with the exception above.

The cause: a missing WithReference.

In Aspire, WithReference(api) is what injects the API's actual endpoint URLs into the referencing project as environment variables (things like services__api__https__0). YARP (and the .NET service discovery middleware underneath it) uses those environment variables to resolve a logical name like api into a real https://localhost:xxxxx address.

Without WithReference, none of those environment variables exist. So when YARP tries to resolve api, there's nothing to resolve it to, and .NET falls back to treating api as a literal hostname. DNS obviously has no idea what api is, and you get a socket exception dressed up as an HTTP error three layers removed from the actual problem.

Remark: this is easy to miss because everything else about the wiring looks correct. WaitFor(api) is there, so startup ordering is fine. WithExternalHttpEndpoints() is there. The project reference at the C# level compiles. It's specifically the environment-variable injection that's missing, and nothing in the dashboard flags that for you.

The fix

One line:

var proxy = builder.AddProject<Projects.Demo_Proxy>("proxy")
    .WithReference(api)
    .WaitFor(api)
    .WithExternalHttpEndpoints();

Add .WithReference(api) before .WaitFor(api), restart the AppHost, and the proxy resolves api correctly. Request goes through, no more socket exception.

That's it.

Takeaway

If you see a SocketException or "host unknown" against a name that looks suspiciously like an Aspire resource name rather than a real hostname (api, db, cache, whatever you called it in AddProject or AddContainer), check WithReference before you check anything else. Service discovery in Aspire is entirely opt-in per reference. Adding a project to the AppHost does not automatically expose its endpoints to the other resources, only WithReference does.

I learnt my lesson…

More information

Popular posts from this blog

Podman– Command execution failed with exit code 125

After updating WSL on one of the developer machines, Podman failed to work. When we took a look through Podman Desktop, we noticed that Podman had stopped running and returned the following error message: Error: Command execution failed with exit code 125 Here are the steps we tried to fix the issue: We started by running podman info to get some extra details on what could be wrong: >podman info OS: windows/amd64 provider: wsl version: 5.3.1 Cannot connect to Podman. Please verify your connection to the Linux system using `podman system connection list`, or try `podman machine init` and `podman machine start` to manage a new Linux VM Error: unable to connect to Podman socket: failed to connect: dial tcp 127.0.0.1:2655: connectex: No connection could be made because the target machine actively refused it. That makes sense as the podman VM was not running. Let’s check the VM: >podman machine list NAME         ...

Cache stampede: when our cache turned against us

While investigating some performance issues, we ran into an ASP.NET Core API that cached a fairly expensive aggregation query for 60 seconds. Under normal load, that was fine: one request rebuilds the cache, everyone else reads from it. Under peak load, dozens of requests would arrive in that same expiry window, all see a cache miss, and all fire the same expensive query in parallel. The database didn't like that. That was the moment when our caching layer stopped helping and started hurting. A burst of requests comes in at the same time, all miss the cache, and all go hammer the database or the downstream API at once. That's a cache stampede . The cache was supposed to protect our backend, and for a few hundred milliseconds it did the opposite. Why this happens IMemoryCache.GetOrCreate (and its async sibling) looks like it protects you, but it doesn't add any locking on its own. Look at the naive version: public async Task<Report> GetReportAsync(string key) ...

A complex system designed from scratch never works

A few years ago, I worked as an architect on a big mainframe rewrite. I still count it as one of my failures. Not because the technology was wrong, but because I couldn't convince the management team to simplify the approach. Years later, the organization is still struggling to get the new system up and running. I left the project at the time, because I couldn't put my name behind an approach that would take very long and cost a lot of money without a working system to show for it along the way. Gall’s Law That memory keeps coming back to me, because it's a textbook case of Gall's Law playing out in real life. Gall's Law , from John Gall's Systemantics , states it plainly: A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch never works, and it cannot be patched to make it work. You have to start over with a simple system that works. What does that mean in practice,...