When the Taxonomy Itself Needs to Scale
Shopify's product classification system handles tens of millions of predictions daily. The taxonomy behind it spans over 10,000 categories and more than 2,000 attributes. But a taxonomy is not static infrastructure—commerce evolves, and the structure that organizes it must evolve too.
The challenge is that Shopify's platform serves over 875 million annual buyers, and the long tail of products merchants list is constantly expanding. New product types, emerging technology categories, and shifting market demands mean the taxonomy needs continuous attention. Manual curation cannot keep pace with that volume, and the alternative—a taxonomy that lags behind what merchants actually sell—has direct costs: reduced discoverability, weaker search results, and less effective filtering for customers.
Three Structural Problems with Manual Curation
Keeping a global product taxonomy current surfaces three distinct issues that manual processes struggle to solve at scale.
Volume and Speed
Every new product type or seasonal trend potentially requires taxonomy updates. Rapidly emerging categories—smart home devices, sustainable products, remote work equipment—carry entirely new attribute sets. Smart home devices need connectivity types, power requirements, and compatibility specifications that simply didn't exist in older category structures.
Domain Expertise Across Verticals
Designing sensible taxonomy requires deep knowledge of each product domain. The distinctions between types of guitar pickups, the correct hierarchy for industrial equipment, or the right attributes for skincare demand specialization across dozens of verticals. No centralized taxonomy team can hold that expertise for every category merchants sell.
Consistency Over Time
Organic growth breeds inconsistency. Similar concepts surface with different representations across categories, naming conventions drift, and merchant-facing categorizations stop matching what customers expect. These inconsistencies compound, degrading both the merchant listing experience and the quality of downstream classification.
From Reactive Manual Review to Proactive AI Agents
The traditional taxonomy workflow was inherently reactive: domain experts analyzed data, identified gaps, and proposed changes only after merchants began listing products that didn't fit. Quality was high, but the process moved too slowly for modern commerce.
The shift came from recognizing that language models could augment—not replace—human expertise. Two complementary analysis modes were needed. One examines the taxonomy itself: gaps in hierarchies, missing attribute relationships, naming inconsistencies. The other grounds analysis in real product data: how merchants describe their products, and what attributes would actually help customers make decisions.
The resulting system combines these modes through specialized agents, producing insights neither approach would find alone.
System Architecture and Analysis Pipeline
The system follows a multi-stage pipeline that combines specialized analysis, synthesis, and automated quality assurance. A critical foundation is enabling agents to interact with the taxonomy itself—searching related categories, examining hierarchies, and validating potential conflicts before changes are proposed. Contextual analysis matters: an agent examining guitar categories can explore the entire musical instruments hierarchy and spot patterns that inform better structural decisions.
The Agent Roles
Four distinct responsibilities are handled by different components within the pipeline:
- Structural analysis examines the taxonomy's logical consistency and completeness, operating purely on the structure to catch naming inconsistencies and organizational gaps.
- Product-driven analysis integrates real merchant data—product titles, descriptions, merchant-defined categories—to find mismatches between how merchants talk about products and the taxonomy's representation of them.
- Intelligent synthesis merges outputs from both analyses, resolving conflicts, eliminating redundancies, and often combining complementary insights into a single improvement.
- Equivalence detection addresses a fundamental commerce tension: merchant flexibility versus platform intelligence. It identifies when a specific category equals a broader category filtered by attribute values. "Women's Golf Shoes" may be equivalent to "Athletic Shoes" with
Activity Type = GolfandGender = Women. Merchants can organize catalogs however suits their business, while platform systems still understand the underlying product relationships.
Automated Quality Assurance Before Human Review
The final stage deploys specialized AI "judges" that evaluate proposed changes before any human review occurs. These judges apply domain expertise and taxonomy design principles to filter suggestions.
Judges are specialized by change type—adding attributes requires different evaluation than creating category hierarchies or modifying existing structures. They are also specialized by vertical: an electronics judge applies technical requirements specific to that industry, while a musical instruments judge brings different expertise. This layering ensures that technical requirements, business rules, and domain knowledge are all weighed before a proposal reaches human curators.
Measured outcomes across taxonomy work
The move from manual curation to AI-assisted evolution has produced concrete improvements in speed, consistency, and coverage. The system's parallel processing lets it analyze whole taxonomy branches at once, turning what used to take weeks into a much shorter review cycle. Where a human expert might work through a few categories per day, the agents evaluate hundreds of categories, checking both structural soundness and alignment with actual merchant product data.
This matters most for emerging categories. When new product types gain traction, the platform can now propose comprehensive taxonomy updates quickly instead of layering on reactive patches that accumulate technical debt.
What the multi-agent design adds
Combining structural analysis with product data surfaces improvements neither method would find alone. Structural checks maintain logical hierarchy and consistency, while product-driven analysis ensures categories and attributes match how merchants actually describe their goods. An automated quality-assurance layer catches potential issues before human review, which has cut down the number of iteration cycles between an initial proposal and final implementation.
One illustrative case involved mobile phone accessories. The product analysis agent noticed merchants frequently advertise "MagSafe support" for chargers, cases, and wallets. It proposed adding a MagSafe compatible boolean attribute to let customers filter for those products. The electronics judge then checked that no duplicate attribute already existed, confirmed the boolean type fit, and recognized that although MagSafe is brand-specific, it works as a legitimate technical standard much like Bluetooth or Qi charging. The judge approved the attribute at 93% confidence, noting it would improve customer filtering for MagSafe-ready accessories.

Proactive maintenance at scale
The larger shift is away from responding to merchant complaints or platform limitations. The system can now proactively spot taxonomy gaps before they start affecting the merchant or customer experience. Because it reasons over the entire taxonomy, it can make cross-category improvements that stay globally consistent, avoiding the fragmentation that comes from solving issues in isolation.
To validate the approach, the team applied this AI-powered method specifically to Electronics > Communications > Telephony (dubbed "Telephony AI" in the analysis) and compared it against the previous manual expansion process as a proof of concept:

Where the system goes next
Several threads are being pulled for future work, all aimed at tightening the loop between taxonomy evolution and the rest of Shopify's classification pipeline.
- Stronger agent reasoning: Newer language models and improved reasoning capabilities could give the analysis agents a better feel for product relationships, catch subtler inconsistencies, and synthesize conflicting insights across approaches.
- Deeper domain coverage: The specialized judge agents are being extended so they can handle more complex categories and emerging commerce trends with more precision.
- Cross-language and regional variation: As Shopify grows globally, the system needs to understand how categorization and attribute relevance change across markets and cultures, allowing regional customization without losing global consistency.
- Tighter classification integration: The eventual goal is a continuous loop: patterns from classification and merchant feedback inform taxonomy priorities, and taxonomy changes immediately feed back into better classification accuracy and acceptance rates.
A shift in how taxonomy work happens
The system marks a fundamental shift from manual, reactive upkeep to proactive, AI-driven evolution. It blends multiple analysis types with automated QA and human oversight, creating something that can scale with modern commerce complexity while retaining the quality merchants and customers rely on.
This is also a demonstration of AI agents augmenting—not replacing—human expertise. The taxonomy team now spends its time on high-level strategic decisions, while the agents handle the broad analysis and quality assurance underneath. The result is a taxonomy that evolves systematically rather than in fits and starts.



