Skip to main content

CosmosDB - Cleanup items automatically

I'm currently attending NDC Oslo. During one of the sessions, this one to be exact, the speaker shared a nice CosmosDB feature; Time to Live.

Some of the typical questions, you ask yourself when building applications are:

  • Should I use hard or soft deletes?
  • Are there any regulations on how long I can keep this data(e.g. GDPR)?
  • How can I keep the amount of data under control?

In the talk, the example they mention is weather data that is collected through various services. This data captured and used to feed an artificial intelligence algorithm to predict the most fuel-efficient way to operate a vessel in the North and Baltic sea.

After a prediction is done, the weather data is no longer necessary and can safely be deleted. Thanks to the Time to Live(TTL) feature in CosmosDB implementing this requirement is really easy.

You have to 2 ways to control the TTL value:

  • At the container level
  • At the item level

Control TTL at the container level

  1. Open the Data Explorer pane for your Azure CosmosDB account.
  2. Select an existing container, expand the Settings tab and modify the following values:

    • Under Setting find, Time to Live.

    • Change the TTL value to On with a value specified in seconds.

    • Select Save to save the changes.

If you prefer to do it through code:

 

Control TTL at the item level

To enable this at the item, we need to introduce an extra ‘ttl’ field on the item:

Is this for free?

One question you probably ask yourself, does this feature comes with a cost? Unfortunately the answer is yes. Deletion of expired items is a background task that consumes left-over Request Units, that is Request Units that haven't been consumed by user requests. Even after the TTL has expired, if the container is overloaded with requests and if there aren't enough RU's available, the data deletion is delayed (although the data will no longer be returned by any queries).

If you want to learn more about this feature, have a look at the documentation here.

Popular posts from this blog

Podman– Command execution failed with exit code 125

After updating WSL on one of the developer machines, Podman failed to work. When we took a look through Podman Desktop, we noticed that Podman had stopped running and returned the following error message: Error: Command execution failed with exit code 125 Here are the steps we tried to fix the issue: We started by running podman info to get some extra details on what could be wrong: >podman info OS: windows/amd64 provider: wsl version: 5.3.1 Cannot connect to Podman. Please verify your connection to the Linux system using `podman system connection list`, or try `podman machine init` and `podman machine start` to manage a new Linux VM Error: unable to connect to Podman socket: failed to connect: dial tcp 127.0.0.1:2655: connectex: No connection could be made because the target machine actively refused it. That makes sense as the podman VM was not running. Let’s check the VM: >podman machine list NAME         ...

Azure DevOps/ GitHub emoji

I’m really bad at remembering emoji’s. So here is cheat sheet with all emoji’s that can be used in tools that support the github emoji markdown markup: All credits go to rcaviers who created this list.

Cache stampede: when our cache turned against us

While investigating some performance issues, we ran into an ASP.NET Core API that cached a fairly expensive aggregation query for 60 seconds. Under normal load, that was fine: one request rebuilds the cache, everyone else reads from it. Under peak load, dozens of requests would arrive in that same expiry window, all see a cache miss, and all fire the same expensive query in parallel. The database didn't like that. That was the moment when our caching layer stopped helping and started hurting. A burst of requests comes in at the same time, all miss the cache, and all go hammer the database or the downstream API at once. That's a cache stampede . The cache was supposed to protect our backend, and for a few hundred milliseconds it did the opposite. Why this happens IMemoryCache.GetOrCreate (and its async sibling) looks like it protects you, but it doesn't add any locking on its own. Look at the naive version: public async Task<Report> GetReportAsync(string key) ...