Algorithmic Responsibility at Spotify: Assessments in Practice
Spotify’s algorithmic systems serve hundreds of millions of listeners and creators, powering personalized music and podcast recommendations. The company has built a framework for algorithmic responsibility that combines centralized research and governance with distributed efforts across product teams. The goal is to evaluate and mitigate potential harms while encouraging recommendation diversity and creating new opportunities for users and creators to connect.
The approach spans three pillars:
- Research: Combining methods across disciplines with product-focused case studies to support equitable algorithmic outcomes.
- Product and tech impact: Working with product teams to translate research into practical changes, supported by organization-wide education and coordination.
- External collaboration: Sharing findings and partnering with researchers, practitioners, and civil society groups.
From Research to Practice
Published research — on topics ranging from gender representation in streaming to voice interface accessibility and the discoverability of underserved podcasts — informs how Spotify evaluates its recommendation systems. Research alone, however, is insufficient. When ML fairness toolkits emerged in the algorithmic impact research community, Spotify's researchers and product teams evaluated their applicability to real-world recommendation systems. That collaboration revealed that sustained effort is necessary to make these toolkits practical, and they often required significant improvement.
Translating research into working tools requires domain expertise. Spotify's algorithmic responsibility researchers work alongside editorial teams and domain experts to address the specific challenges of music, podcasting, and other content types. Cross-functional collaboration with legal and privacy teams helps ensure that impact assessments meet privacy standards and comply with applicable regulations.
Scaling the Effort
Expanding algorithmic responsibility across the organization has involved several strategic moves:
- Building a dedicated central team of qualitative and quantitative researchers, data scientists, and NLP specialists to develop methods and track company-wide impact.
- Creating an ecosystem of responsible practices by launching targeted collaborations with product, machine learning, and search teams.
- Introducing governance structures, central guidance, and best practices for personalization, data usage, and content recommendation.
How Impact Assessments Work
Algorithmic impact assessments (AIAs) are internal audits used to evaluate music and podcasting systems. They serve as learning tools rather than formal code audits, helping product owners create roadmaps that account for potential effects on listeners and creators. AIAs highlight where deeper investigation may be needed, whether that involves product-specific guidance, technical assessment, or external review. Like the products themselves, the assessment process requires iteration and refinement over time.
Assessments also increase accountability. Since audit results are shared with stakeholders across the organization and work is tracked on product roadmaps, product owners are more likely to prioritize responsible design. The process includes appointing formal product partners who work directly with the core algorithmic responsibility team.
Within less than a year, Spotify educated a significant number of employees and assessed more than one hundred systems involved in personalization and recommendation. Results include formal work streams that improved recommendations and search, better data availability for tracking platform impact, and reductions in data usage. Some teams implemented additional safety mechanisms to avoid unintended amplification, and efforts to address dangerous or sensitive content were informed by the assessments.
Lessons from Implementation
Operationalizing AIAs across a large organization has surfaced several lessons that extend beyond Spotify's specific context.
Principles Need Continual Investment
Turning responsible ML principles into day-to-day engineering practice requires practical structures, clear guidance, and playbooks that address real-world challenges. This work resists completion: as products and models evolve, so does the need to evaluate them at scale. Some cases will demand tools beyond what now exists in the research or industry community; others will require automating processes to reduce the burden on product teams.
Responsibility Cannot Be Centralized
Algorithmic responsibility cannot be the job of a single team. Sustained improvement requires technical and organizational change, which means every team must understand its roles and priorities. Spotify has focused on organizing people and processes — not just machine learning systems — to embed responsible practices as business as usual for hundreds of employees.
Systems Require a Holistic View
User experiences depend on many interdependent systems, algorithms, and behaviors working in concert. Changing a single system may not produce the intended effect. Coordination across teams, business functions, and organizational levels is essential to understand critical dependencies and develop tracking methods that span system types.
Internal Audits Need External Perspectives
Internal processes like AIAs are complemented by external auditing and collaboration with academic, industry, and civil society groups. Spotify works with its Diversity, Inclusion & Belonging team, editorial teams, and the Safety Advisory Council, as well as researchers from disciplines including social work, HCI, information science, computer science, privacy, law, and policy. Cross-disciplinary translation is challenging and time-intensive, but the investment is essential.
Solutions Require Cross-Sector Translation
The timelines of product development and research insights frequently diverge, so expectations must account for the fact that algorithmic impact problems are not always solvable through technical means alone. Effective assessment combines domain research on what matters in music and podcasting with internal research on how to create structure and support within existing organizational processes. Spotify is investing in data, infrastructure, and research to improve impact tracking, while sharing best practices across academia and industry through conferences, workshops, and formal collaborations.
What the front-line teams taught us
Running the assessment itself surfaced a recurring operational truth: the people closest to a product know where its risks actually live. Spotify's central AI governance function, working with product, engineering, design, data science, and user research teams, found that generic checklists only get you so far. The useful signal came from structured workshops where each group mapped their specific feature's risk surface — and from honest conversations about what could go wrong, not just what was built to go right.
Three patterns stood out across teams:
- Risk identification is a team sport. Engineers know the model's failure modes; designers know where the UX can mislead; researchers know which user segments are underrepresented. No single role can pre-empt all harms.
- Impact assessments are iterative. The first pass at a risk register was never the last. New edge cases emerged in testing, in beta, and after launch — and the assessment had to be treated as a living document, not a launch gate artifact.
- Remediation requires ownership. The most effective mitigations came from teams that were empowered to change product behavior, not just add a flag or a log line. Where ownership was ambiguous, risks lingered.
The hard cases: aggregation, proxies, and contrafactuals
A recurring analytical challenge was deciding where to draw the boundary of an assessment. A recommendation model, for example, might be assessed on its own — but its impact is shaped by upstream ranking, downstream content policies, and the user's own history. Aggregating risk across a complex system tends to dilute it; breaking it down too finely misses the compounding effects. Spotify's teams found that assessing at the level of a user-visible experience, rather than a single model artifact, usually captured the most meaningful harms.
Another lesson involved proxies. Several product metrics that seemed like useful risk indicators turned out to proxy for something else — for instance, engagement time can signal immersion or over-reliance. Teams learned to state their assumptions about what a metric means before using it as a proxy for harm, and to validate those assumptions against qualitative user feedback.
Finally, a practical note about contrafactuals. In an assessment, it is tempting to ask "would this harm happen without the algorithm?" — but in a product where the algorithm has always been present, that counterfactual often has no clean answer. Spotify's approach was to compare against a baseline of the product's prior behavior, not an idealized non-algorithmic system.
Where the practice is heading
Algorithmic impact work at Spotify remains iterative and collaborative. The governance team continues to refresh its methodology as products, content, and regulations evolve; the goal is to keep the practice informed by real incidents and close to the teams building the systems. The company's public research on this topic is collected at research.atspotify.com/algorithmic-responsibility, and the team acknowledges the ongoing guidance of the Spotify Safety Advisory Council in shaping these efforts.



