Lifecycle Stages for Infrastructure at Slack

Around 2020, Slack’s Cloud Engineering organization — later split into several more focused teams — oversaw Infrastructure-as-a-Service providers, compute environments (Chef, Kubernetes) and a wide range of related systems. The breadth of that remit became a problem: multiple technologies overlapped in purpose, features were reimplemented instead of built once, and retiring older systems was hard because migrations were ill-defined.

Worse, communication with internal users was inconsistent. What a peer heard about support or deprecation depended on which engineer they asked — if they got an answer at all. There was no single model for describing what a team could expect from a given technology, which made it difficult to recommend one option over another.

What a Lifecycle Needed to Solve

Slack designed a lifecycle model with three explicit goals:

  • Standardization. Status and expectations had to be uniform, regardless of which engineer was consulted or which technology was in question. Two systems slated for retirement should be handled identically, with executable plans for getting there.
  • Clarity. Advice about technology choice was often opinion-based. The goal was to present factual status — support level, future trajectory, production readiness — so peers could make their own decisions rather than weigh subjective judgments.
  • Autonomy and control. A lifecycle gives the owning team a mechanism to retire systems that are superseded by better options — but retirement only happens when dependent teams migrate. The lifecycle had to give those teams enough forewarning to plan around their own constraints.

What Gets Tracked Per Stage

Each lifecycle stage defines concrete expectations for:

  • Acceptance of new customers — ranging from “not yet” to “prohibited”
  • Feature requests — whether they’re accepted and how they’re prioritized
  • Bug reports — the same, applied to defects
  • Security and compliance reports — always accepted, though the resolution may be shutdown rather than a fix
  • Documentation quality — how current and accurate it can be expected to be
  • Progression rules — for example, an Active system never goes straight to Retired

Assignment is made by team consensus: is the system feature-complete and bug-free enough for broad use, can the team support it at the corresponding level, and — for later stages — is there a feature-complete replacement with sufficient migration communication? Movement into later stages includes dialogue with actual users, to ensure that replacements the team considers viable are ones those users can adopt.

Stage 1: Alpha

Definition. An experiment or pre-prototype. Likely to be discarded, though a “v2” may become Beta. Hack day projects and scoping work are typical examples.

Availability. General use is strongly discouraged.

Support. None. Bug reports and feature requests may be summarily declined, except for security/compliance bugs.

Roadmap. Advances to Beta or Retired, depending on what the experiment shows.

Stage 2: Beta

Definition. A technology built to become Active, but not yet ready for broad consumption. A prototype.

Availability. Self-service production use is strongly discouraged; limited production use is allowed only in close consultation with the owning team. Non-production use is fine for early adopters.

Support. Provided to gather feedback that advances the project. Documentation is sparse at best; Slack discussions are the most reliable source of information. All non-security/compliance requests are low priority or scheduled on the roadmap.

Roadmap. On track for Active, but may instead go to Maintenance — rarely straight to Deprecated.

Stage 3: Active

Definition. Production grade, with staffed support. These are the flagship technologies the team hopes everyone will adopt.

Availability. Available for production use.

Support. Resources are reserved for bug reports and feature requests, prioritized per normal procedures. Most Active technologies also have reserved time for hands-on support and consultation — not doing operational tasks for users, but working alongside them to meet their objectives.

Roadmap. Transitions to Maintenance eventually; very rarely goes straight to Deprecated, but never directly to Retired.

Stage 4: Maintenance

Definition. Older technologies whose replacement is late in Beta or early in Active. This is the typical transition point.

Availability. Still available for production use, but new use cases are discouraged, not prohibited. Active (or sometimes Beta) alternatives are preferred.

Support. Resources cover bug reports but not feature requests. Feature requests may be summarily declined. Hands-on consultation is generally not provided, and may be scoped to designated office hours depending on the maturity of alternatives.

Roadmap. May move to Deprecated, or alternatively be revived to Active if resources and priorities allow.

Stage 5: Deprecated

Definition. A technology with a feature-complete Active alternative. These often carry external hard stop dates, such as an operating system's end-of-life.

Availability. Existing users should migrate as quickly as possible; new use cases are prohibited.

Support. Minimal. Documentation may be outdated or wrong. Bug reports other than security/compliance are low priority, and feature requests may be summarily declined.

Roadmap. Remains in Deprecated for long enough that reasonable users can migrate without haste — typically at least two full quarters after the deprecation date. After that, retirement is at the team’s discretion.

Stage 6: Retired

Definition. No longer supported and not intended for use. Remnants may remain and might even function, but they can fail at any moment — in which case they are deleted, not repaired.

Availability. Not available for any purpose. New use cases are forbidden; existing ones should be migrated immediately.

Support. None. Even security/compliance requests are answered by deleting the related infrastructure.

Roadmap. End of the line.

Lifecycle in practice: three generations of compute

The Technology Lifecycle framework was built as a communications and planning tool, and its value became clear as we steered three successive iterations of our platform-ized Compute offering. Mapping resource allocation against each technology’s lifecycle stage revealed both over-investment and under-investment relative to where a given platform was headed.

BuiltBySlack (BBS)

February 2017 saw the launch of our first internal platform, BuiltBySlack ("BBS"): a rudimentary framework of parameterized Terraform modules and Chef cookbooks that deployed and ran services bundled in an internally-developed packaging format. While basic, BBS proved there was both a use case and an appetite for a platform offering, and it carried some critical infrastructure along the way.

Bedrock Classic

By May 2018, the abstraction and standardization benefits of BBS were clear, prompting work on a Kubernetes-based offering called "Bedrock". The first iteration, Bedrock Classic, reused much of the same infrastructure as our current generation container platform but pushed more operational responsibility onto service owners: they wrote shell scripts to build images, pushed them to registries, and authored their own Kubernetes resources to run those images.

Bedrock YAML

The gap between what operators had to manage and what they wanted to abstract grew, leading to the next generation, Bedrock YAML, which automated more of the build-and-deploy path.

Applying the lifecycle

When the Technology Lifecycle was introduced in March 2020, it immediately gave us a shared vocabulary. Bedrock Classic was labeled Active in the first draft, BBS was Maintenance, and Bedrock YAML was Beta. Those labels turned into decisions:

  • July 2020: With feature parity agreed upon by both the platform team and remaining BBS customers, BBS was moved to Deprecated.
  • April 2021: Based on Beta customer experience and internal production readiness assessments, the Cloud Engineering team promoted Bedrock YAML to Active, signaling full production readiness.
  • July 2021: All BBS customers had migrated to Bedrock, and BBS—along with its proprietary packaging format—was moved to Retired.
  • July 2022: After a long stretch of maintaining feature parity between the two Bedrock iterations, Bedrock Classic was moved to Deprecated, communicating to its customers that migration to Bedrock YAML was the expected path.

The payoff

Across both this platform journey and a more regimented process for products like Ubuntu releases, a concrete, documented Lifecycle and Maturity Model made planning and communication straightforward. It also made it far easier to stop investing in technologies that no longer fit the future we were building toward. Cutting spend on systems marked for deprecation allowed a small, agile team to focus on high-impact work serving a much larger engineering organization.