Home/Engineering/A Guide to Running an Engineering Program - Shopify
Engineering
A Guide to Running an Engineering Program - Shopify
Shopify has a playbook for huge engineering programs in the past that cross over multiple areas of complexity on the platform, and we’d like to share it with you.
SE
Shopify EngineeringShopify Engineering
·July 21, 2021
Running Large Engineering Programs at Scale
After 15 years of building commerce tooling and scaling architecture to support roughly 8% of the world's adult population, Shopify's platform complexity has grown significantly. The company has taken on sprawling engineering programs—from the Storefront Renderer rewrite to BFCM performance work, platform-wide capacity planning, resiliency testing, and static typing tooling. A documented playbook emerged from those efforts, centered on a few repeatable structures that keep multi-team programs aligned and moving.
Program Definition: The North Star
Every program starts with a clearly articulated definition of done. That clarity keeps the team pointed in the same direction and produces assets that enable company-wide alignment, contextual status updates, risk mitigation, and decision-making. The Program Plan—covering duration, scope, staffing, and outcomes—must be agreed upon by all stakeholders, typically Directors and VPs in the affected areas. They inherit any technical debt or decisions made along the way, so their buy-in matters from day one. Together with the Program Lead(s), these stakeholders form the Program Steering Committee and frame out the following components:
Problem Statement
Articulate exactly what the issue is, with enough specificity to win executive buy-in. Useful questions: What can users or the company do after this goal is achieved? Why solve this now? What are we deprioritizing to make room, and is that tradeoff correct?
Program Objectives
Objectives serve as a motivational drumbeat and a negotiation lever when resources tighten or competing programs gain traction. The same framing questions apply: what becomes possible when we finish, what does the company gain, and what are we choosing not to do?
Guiding Principles
Spend real time here—guiding principles become the tool for making tradeoff decisions and forcing constraints on proposed solutions.
Definition of Done
The program-level definition of done is what the stakeholder group is accountable for; each workstream has its own, owned by contributors. Program Leads manage both. A complete definition includes a checklist per team, a performance baseline, a resiliency plan with a gameday, and both internal and external documentation. Setting these expectations upfront lets teams parallelize confidently.
Top Risks and Mitigation
Identify the most realistic blockers to reaching the definition of done and what you'll do about them. This list is a living artifact—mitigate one risk and another typically emerges. Risks are often technical but can be resource-based if staffing doesn't arrive on schedule.
Path to Done
With a clear start and end state, define the timeline and how you'll traverse it. Scope and staffing act as levers, adjusted holistically alongside other company initiatives to avoid over-optimizing one program at the expense of the whole. The full set of assets—problem statement, objectives, principles, definition of done, risks, and path—gives stakeholders the context to decide staffing duration and program execution.
Execution itself remains the most nebulous phase of any large, high-stakes program. As Eisenhower put it, "Plans are worthless, but planning is indispensable." The plan's value shows up in implementation.
Execution Rhythm
At Shopify R&D, the company operates on six-week cycles. The definition of done acts as the program's primary reinforcing loop—until the North Star is reached, each cycle factors in unachieved goals from the previous cycle, newly surfaced risks, and the expected targets from the Path to Done. A cycle opens with clear goals and moves into execution, or as Shopify says, GSD (get shit done).
Six week cycle structure
The mechanics of aligning on a Program Plan are usually less challenging than sustaining momentum. What holds a large program together is the way teams work within each cycle and the rituals that keep a 200-person, 92-project effort on track. Those rituals break into three cadences: weekly, every-six-weeks cycle rituals, and ad hoc checkpoints.
Making Progress Visible Every Week
Program communication works best when it runs on a predictable cadence. Shopify’s program leads use a set of weekly rituals to keep stakeholders informed, clear blockers, and keep the whole team aligned on goals. Each ritual has a clear owner, a recurring calendar slot, and a concrete deliverable that feeds into the next ritual.
The Company Wide Program Update
Early each week, program leads draft and review a company-wide status update. The update is written against a shared template that mirrors the format and language of the cycle goals from the Program Stakeholder Review, so stakeholders can read the plan the same way each cycle. The document tracks progress across every active workstream and is also a place to surface risks, blockers, concerns, and celebrations.
The drafting and review are scheduled as two separate recurring calendar events to ensure the communication goes out on a predictable schedule. During the day, program leads collaborate on the document, flagging tripwires that might require ad hoc outreach. If the review meeting finds nothing missing, it is cancelled and the time is given back. If gaps remain, the team uses the meeting to wrap the update together.
Deliverable: A weekly email to stakeholders with status against the current cycle goals.
Status Check-Ins Between Program and Project Leads
Project leads finish their individual project updates by Friday afternoon, feeding into Monday’s company-wide update. A recurring check-in between program and project leads early in the week provides dedicated space for cycle changes, program announcements, or housekeeping. If the company update was completed cleanly, the sync is cancelled for everyone except the few leads whose updates were missing.
Deliverable: An accurate picture of each project’s status and its likelihood of hitting its cycle goal, which directly informs the company-wide update.
Escalation Triage
Escalations are tracked and triaged at least weekly, and often ad hoc, through the Weekly Program Lead Sync. Program leads log escalations as they arise into a GitHub project board, using tags to sort by urgency and required timing. They also monitor technical designs, project updates, and team demos to catch issues before they become blockers.
An escalation’s outcome is often a decision. Once key stakeholders align on the decision, it is recorded in the decision log. The triage process delegates ownership and ensures a solution is prioritized within the program team.
Deliverable: Clear ownership of each escalation. Aggregate blockers and solutions surface as highlights in the Program Stakeholder Review.
Proactive Risk Triage
Risks get the same weekly and ad hoc attention as escalations, but with a forward-looking lens. The planning spreadsheet holds a ranking formula that prioritizes which risks need mitigation first. For each risk, the sheet identifies where it sits in the program and which lead owns the mitigation strategy, along with a last-updated date for the mitigation status.
That last-updated date matters: it lets program leads check in without repeatedly pressuring teams for an immediate update. Once a mitigation plan is in place, the team updates the sheet with the plan and its implementation status. Only when the plan is implemented does the risk ranking change and the risk actually get retired. Collaboration happens through comments in the spreadsheet and in Slack channels, where program leads can celebrate wins and reinforce momentum.
Deliverable: Delegated ownership of mitigation plans. Top risks and their mitigation status aggregate into the Program Stakeholder Review as highlights.
Program Lead Sync
Program leads meet weekly to strategize, collaborate, and divide up work. The agenda starts from a few standing bullets to keep the conversation focused on partnership rather than coordination:
Real Talk: What is top of mind and what is keeping you up at night.
Demo Plan: What messaging to deliver when the full program team gathers for demos.
Divide and Conquer: Which meetings can be dropped to reduce redundant coverage.
Risk Review: Current top risks and how mitigation plans are shaping up.
Additional agenda items are added throughout the week, usually escalations that could affect program velocity or a project’s scope.
Deliverable: A communication and messaging plan for the weekly demo, with current risk levels and mitigation status based on time passed and any new information or tooling changes.
Weekly Demos as the Team Gathering Point
Weekly demos fill two roles depending on where the program sits in its six-week cycle. In weeks one through five, demos are a space for the team to share progress and celebrate contributions to the goals. In week six, they showcase the planned goal progress to stakeholders. The session is scheduled for the end of the day on Fridays.
Preparation happens in two tracked: planning the demos themselves and planning their facilitation. Project leads can sign up to demo, with a call going out about two days in advance. They mark their intent on the planning spreadsheet. Meanwhile, program leads and domain leadership figure out facilitation: announcements, timed messaging, and the order of demos. The tone stays light, with room for jokes, music, and a shared team vibe while demoing work and answering questions.
Deliverable: A recorded session available company-wide for review and follow-up questions, linked from the weekly Company Wide Program Update.
Cycle Rituals: The Six-Week Cadence That Keeps Programs on Track
Cycle Kick Off
Every new cycle begins on day one of week one with a full team sync-up. The entire program team is invited. The goal is straightforward: align on what we’re working toward, share progress, and onboard any new team members or workstreams. It also gives everyone visibility into parallel projects, helping the team anticipate changes and collaborate on shared patterns early.
The session is kept short and energetic. It’s a live presentation covering the cycle’s investment plan, overall program progress, and the biggest risk areas expected over the next six weeks. Recurring calendar items handle the scheduling, and reminders about operational details—like the definition of done or office hours—are raised here so support structures stay front of mind.
Mid-Cycle Goal Iteration
Goals aren’t always realistic when set; that only becomes clear once work begins. Between weeks one and three, Project Leads are empowered to change their cycle goal—no more than once per project—provided they explain why and propose a new goal attainable in the remaining time. This flexibility keeps the plan honest without letting it drift.
Leads share these evolutions over Slack, making sure the impact on subsequent cycles in the program plan is understood. Week three brings a reminder paired with Goal Setting Office Hours, so no one misses the window to adjust.
Goal Setting Office Hours
Week three is reserved for reviewing current cycle goals. Weeks four and five shift focus to the next cycle’s goals. This deliberate scheduling means the next cycle’s plan is built intentionally rather than rushed once the Program Stakeholder meeting is booked—leaving the team aligned well before the week one kick off.
Office hours are set up through a recurring calendar event on the shared program calendar, with a sign-up sheet for individuals to claim time. It’s not a heavily used process, but the availability and guaranteed time slot give Leads confidence that support exists if needed. The Program Touch Base ritual often catches risks and velocity changes before office hours become necessary, though the team hasn’t yet determined whether office hours could be removed entirely.
Cycle Report Card
At week five, Slack nudges Leads to begin reflecting on the cycle’s wins. Over the next week, nominations come in highlighting the team’s best achievements measured against company values—collaboration, merchant/partner/developer obsession, and resourcefulness. The result is a templated report card that looks back at what was set out to be done and shows the team the velocity and impact of their work.
This is delivered by Program Leads in a full team sync-up, with Team Leads providing the praise. It’s a moment for gratitude and celebration, and it doubles as a way to demonstrate alignment among Program Leads while helping the entire team understand its collective strengths better. The celebratory section then carries into the next Cycle Kick Off presentation, reflecting on both individual contributions and team collaborations.
Program Lead Retro of the Previous Cycle
Every six weeks—skipping the first cycle—Program Leads run a stop-start-continue workshop to reflect on recent experience and decide what should change. The retro typically lands in week one, after Project Leads have shared their own feedback. Participants consider:
How Program Leads work together as a team.
How Program Leads manage up to Program Stakeholders.
How Project Leads manage up to Program Leads.
What feedback Team Leads are providing.
How program execution is going within each team.
Added questions help steer thinking about what feedback actually matters. The workshop produces lessons that drive ritual changes, reviewed in the Lead Sync starting week two. Program Leads aim to implement and communicate changes to the broader team by the end of week two, leaving four weeks for the changes to take effect before the next retro. The documented summary is made available company wide and included in the Program Stakeholder Company Wide Update.
Project Lead Retro of the Previous Cycle
Project Leads have the option to run their own retro as part of their rituals, held in week six or week one while the cycle is still fresh. Even if a Project Lead decides not to run one, a Program Lead can still request it. The approach is not prescribed beyond Shopify’s general Get Shit Done recommendations—what matters is the outcome, not the format.
An anonymous feedback form is shared by Program Leads in advance of week six, asking the team what to stop, start, and continue, including at the program level, plus an open-ended section to distill lessons learned. These lessons are shared back with all Project Leads, fostering a “first team” mindset as described by Patrick Lencioni: true leaders prioritize supporting their fellow leaders over their direct reports. Teams that go far and fast adopt this mentality because a foundation of trust and understanding makes change, vulnerability, and problem-solving far easier. As the thinking goes, ideas and plans don’t solve problems—teams do.
Program Stakeholder Review
Early in week six, often booked by the VP office, Program Stakeholders review the upcoming cycle’s goals. This is the forum to set expectations, escalate risks, and discuss scope changes in light of other goals and decisions—always framed against both the cycle ahead and the overall program plan.
Program Leads provide a status update on the previous cycle and visual goals for the next one. Leading up to it, the Weekly Sync is used to align on how to use the stakeholders’ time, focusing on the most important discussion points and making the most of the Program Steering Committee’s attention. The deliverable is a presentation covering progress, the remaining plan, open decisions, and escalations needing additional support.
Program Stakeholder Company Wide Update
At least once per cycle, typically at the start of week four, the program sends a company-wide status update. This timing follows the Mid-Cycle Goal Iteration window, clarifying any changes made during that period. Shopify’s internal culture prizes transparency—it’s one of the things that makes the company fast. Publishing program status and its evolution cycle to cycle creates intense collaboration, letting teams spot dependencies and risks early.
Two recurring calendar events—one to draft, one to review—keep the schedule predictable. All Programs use a shared template that mirrors the language and layout of the original program kickoff, so stakeholders can interpret the update without any learning curve. During the day, Program Leads collaborate on the document, flagging areas that need attention or support, updating forecasts, highlighting plan changes, and often celebrating a program that’s on track. The final update reaches stakeholders by email.
Rituals for the Unexpected
Even with a full calendar of recurring ceremonies, the unpredictable elements of an engineering program need their own set of responses. These ad hoc rituals are triggered by judgment and context rather than a schedule. They exist to capture the operational and technical details that can shift a program’s scope or velocity, navigating the uncertainties that aren’t visible in the initial plan.
Assumption Slam Workshop
Large programs sit at the intersection of product, UX, and engineering, and that intersection is fertile ground for miscommunication rooted in unspoken assumptions. The Assumption Slam is a facilitated workshop designed to surface those assumptions early so they don’t become blockers. It is most critical within the first month of kickoff.
In weeks one through three, the program leads run this guided session with a deliberately small group. The facilitator needs enough program context to press the team on the root of each assumption and its downstream impacts. The right participants must be present to challenge plans at the true points of intersection. Each key item identified during the session becomes an action item — either to mitigate a risk, finalize a decision, or launch a deeper investigation.
Program Touch Base
This is the conversational safety net between the program lead and each project lead. It is triggered by instinct as much as by data. Typical triggers include:
A workstream’s goals have been off track for more than a week without communication.
A status moves directly from green to red, skipping yellow.
Team updates are consistently late or missing.
A lead has gone a full cycle without contact.
New information also triggers the touch base — for example, another initiative that changes scope, dependencies, or staffing. The meeting itself follows a light structure with a standard agenda that includes:
Real Talk: What is on the lead’s mind? This keeps the relationship a partnership rather than a status check.
Confidence check: Can the lead deliver this cycle’s goal? What about the full program schedule, including time, staffing, and scope?
Challenges and risks: What obstacles are in the way?
Performance: How is the project shaping up on speed, scale, and resiliency? Are there concerns that could jeopardize the workstream’s definition of done?
What aren’t you doing? Stakeholders often inherit the technical debt of decisions; asking this surfaces items for the roadmap backlog.
The deliverable here is engagement: clarified dependencies, identified risks, and decisions made before they become problems.
Engineering Request for Comments (RFC)
In a single project, technical design documents build consensus about the whole direction. In a program with many interconnected workstreams, the team needs to align quickly on smaller technical areas without revisiting the entire project. The RFC ritual fills that gap, using an asynchronous GitHub-based approach with a documented template and rules of engagement. If alignment is not reached by the deadline and no one has explicitly vetoed the approach, the RFC author makes the final call. The output is a documented architectural decision — fast consensus without the overhead of a full design review.
Performance Testing
Performance is a core attribute of Shopify’s products: it cannot regress, and any improvement is a notable win for the Cycle Report Card. Performance testing is triggered by deploying to production and is run both on individual components and as part of the integrated system.
Teams write a load test for their component by configuring a shop and scripting with Lua, orchestrated through internal tooling called Genghis. Results are validated against the program’s Service Level Indicators. If a component passes in isolation, the team folds it into a full end-to-end test where system complexity surfaces. This happens through async discussion and office hours hosted by the program’s performance team, which documents context and inherits the testing flows and associated shops. Multiple shops and flows are used deliberately: services are tested on the happy path and under maximum complexity to reveal how the system actually behaves.
The immediate deliverable is a feedback loop confirming the service meets performance expectations. Longer term, the performance team gains the ability to run recurring regression tests and stress the system on any dimension they choose.
Engineering program management is still an evolving craft, adapting to the needs of the specific program, the company’s organization, and the management structures of the teams involved. Not every ritual is necessary in every context, but in combination, these scheduled and ad hoc ceremonies keep a 200-person program of 90 projects on track without burning out the team. This approach to program design supported the outcomes Shopify announced at Unite 2021 — custom storefronts, checkout app extensions, and Online Store 2.0.
About a year ago, I was offered a presentation slot at the WeAreDevelopers World Congress in Berlin. I rarely take speaking engagements, especially international ones, but this one arrived at just the right time, the right place, and with the right person – I said yes, on the contingency that Ben Dumke-von der Ehe joins me in the presentation. Ben is an early community hire at Stack Overflow who l