Skip to main content

You don’t have a platform if it doesn’t have self service

Although the concept is not new (it was introduced in the Thoughtworks technology radar in 2017), I see a recent grow in platform teams at my customers. Partially this could probably be explained by the success of the great Team Topologies book that can be found on the bookshelf of almost every IT manager today.  

What are platform teams?

Platform teams are specialized groups within an organization that focus on building and maintaining the foundational technology and infrastructure that other development teams use to create applications. Their primary goal is to provide reusable tools, frameworks, and services that streamline the development process and enable feature teams to focus on delivering business value without worrying about underlying technical complexities.

Platform teams typically handle:

  • Infrastructure Management: Setting up and maintaining cloud services, CI/CD pipelines, monitoring tools, and other foundational infrastructure.
  • Developer Tools: Creating and managing internal tools and libraries that help developers write, test, and deploy code more efficiently.
  • Standardization: Establishing coding standards, best practices, and common frameworks to ensure consistency and quality across the organization’s software projects.
  • Automation: Automating repetitive tasks and processes to reduce manual effort and minimize the risk of human error.

Pitfalls when creating a platform team

The growing number of platforms teams made me suspicious and when taking a look in more detail at some of those teams, I noticed the following things:

  • Command and Control: Instead of being an enabler for other teams in the organization, they command how other teams should work making them a bottleneck for the rest of the organization. A lot of times when I noticed this, the platform team was a ‘traditional’ IT support team in disguise.
  • Overengineering Solutions: The team went overboard and created overly complex tools and frameworks that are difficult for feature teams to use. Striking the right balance between sophistication and usability is crucial.

  • Insufficient Collaboration with Feature Teams: The team operated in isolation and developed solutions that do not address the real pain points of the feature teams. Continuous collaboration and feedback loops are essential to ensure that the platform team's efforts are effective and relevant.

  • Neglecting Documentation and Training: Even the most powerful tools and frameworks are ineffective if the development teams do not know how to use them properly. Comprehensive documentation and training are critical to the successful adoption of platform team deliverables.

  • Lack of Self-Service:  Platform teams require a significant investment in skilled personnel and resources. Automation and self-service are key to create a scalable solution without the platform team becoming overworked and a bottleneck.

This all made me think about the following video from Sam Newman:

In this video Sam talks about his own experience with platform teams and noticed the following pitfalls(some of them similar to my own experience):

  • Not enabling self-service: Feature teams should be able to move fast and be empowered to make decisions and get things done.
  • Not helping people use the tools well: the job the platform team is not to build a platform, it s about enablement.
  • Trying to implement governance through tooling: Forcing people to use your platform isn’t about enablement, it’s about control.

His advices are simple:

  • Trust your people
  • Treat your platform as a product
  • Make your platform optional
  • Provide a paved road experience

Let me end this post with the following quote from the video above:

If you have to ask to a person to get something done it is not a platform

More information

Platform engineering product teams | Technology Radar | Thoughtworks

Team Topologies

The paved road (bartwullems.blogspot.com)

Building platforms–Strike the right balance (bartwullems.blogspot.com)

Popular posts from this blog

Podman– Command execution failed with exit code 125

After updating WSL on one of the developer machines, Podman failed to work. When we took a look through Podman Desktop, we noticed that Podman had stopped running and returned the following error message: Error: Command execution failed with exit code 125 Here are the steps we tried to fix the issue: We started by running podman info to get some extra details on what could be wrong: >podman info OS: windows/amd64 provider: wsl version: 5.3.1 Cannot connect to Podman. Please verify your connection to the Linux system using `podman system connection list`, or try `podman machine init` and `podman machine start` to manage a new Linux VM Error: unable to connect to Podman socket: failed to connect: dial tcp 127.0.0.1:2655: connectex: No connection could be made because the target machine actively refused it. That makes sense as the podman VM was not running. Let’s check the VM: >podman machine list NAME         ...

Cache stampede: when our cache turned against us

While investigating some performance issues, we ran into an ASP.NET Core API that cached a fairly expensive aggregation query for 60 seconds. Under normal load, that was fine: one request rebuilds the cache, everyone else reads from it. Under peak load, dozens of requests would arrive in that same expiry window, all see a cache miss, and all fire the same expensive query in parallel. The database didn't like that. That was the moment when our caching layer stopped helping and started hurting. A burst of requests comes in at the same time, all miss the cache, and all go hammer the database or the downstream API at once. That's a cache stampede . The cache was supposed to protect our backend, and for a few hundred milliseconds it did the opposite. Why this happens IMemoryCache.GetOrCreate (and its async sibling) looks like it protects you, but it doesn't add any locking on its own. Look at the naive version: public async Task<Report> GetReportAsync(string key) ...

A complex system designed from scratch never works

A few years ago, I worked as an architect on a big mainframe rewrite. I still count it as one of my failures. Not because the technology was wrong, but because I couldn't convince the management team to simplify the approach. Years later, the organization is still struggling to get the new system up and running. I left the project at the time, because I couldn't put my name behind an approach that would take very long and cost a lot of money without a working system to show for it along the way. Gall’s Law That memory keeps coming back to me, because it's a textbook case of Gall's Law playing out in real life. Gall's Law , from John Gall's Systemantics , states it plainly: A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch never works, and it cannot be patched to make it work. You have to start over with a simple system that works. What does that mean in practice,...