How the Core Web Vitals Thresholds Were Chosen
Core Web Vitals are a set of field metrics designed to measure key aspects of real-world user experience. Beyond the metrics themselves, each has target thresholds that classify performance as "good", "needs improvement", or "poor". The methodology behind these thresholds is based on research into human perception, real-world data on achievability, and careful consideration of how to classify overall site performance.
The Three Core Metrics
The Core Web Vitals cover three distinct aspects of user experience:
- Largest Contentful Paint (LCP): Measures perceived load speed by marking the point when the page's main content has likely loaded.
- Interaction to Next Paint (INP): Measures responsiveness, quantifying the experience users feel when interacting with the page.
- Cumulative Layout Shift (CLS): Measures visual stability by quantifying unexpected layout shifts of visible content.
Each metric's thresholds categorize performance into "good", "needs improvement", or "poor". To classify a page or site overall, the 75th percentile value of all page views is used. If at least 75% of page views meet the "good" threshold, the site is classified as "good" for that metric. Conversely, if at least 25% of page views fall into the "poor" range, the site is classified as "poor". For example, a 75th percentile LCP of 2 seconds is "good", while 5 seconds is "poor".
Criteria for Setting Thresholds
Several criteria guided the threshold selection process, and these criteria sometimes conflict with one another. The goal was not to find a single perfect threshold, but to select the candidate that best satisfies the criteria overall.
High-quality user experience. The primary goal is optimizing for the user. Where human perception and HCI research exists, the thresholds are aligned with consensus ranges from the literature. However, this research is often expressed as a range rather than a single value, and aggregated Chrome metrics data shows user behavior changes gradually rather than at a fixed point. In cases where no relevant UX research exists—such as with a new metric like CLS—real-world pages meeting different candidate thresholds were evaluated to identify which threshold corresponds to a good experience.
Achievability for existing content. A threshold must be something site owners can realistically meet. While an LCP of zero milliseconds would be ideal, it is not practically achievable due to network and device latency. To verify achievability, thresholds were tested against Chrome User Experience Report (CrUX) data, requiring that at least 10% of origins meet the "good" threshold. The thresholds also needed to be consistently met by well-optimized content, so that good sites are not misclassified due to field data variability. For the "poor" threshold, the default approach was to classify the worst-performing 10-30% of origins as "poor", unless relevant research suggested otherwise.
Device consistency. Mobile and desktop have very different device capabilities and network reliability, which heavily impacts achievability. However, users' expectations of a good or poor experience do not depend on the device. For this reason, the same thresholds are used across both device types, which also keeps the thresholds simpler to understand and avoids the complexity of categorizing devices by form factor, processing power, or network conditions. Because mobile is more constrained and often represents the majority of traffic, the thresholds effectively reflect mobile achievability.
Choosing the 75th Percentile
The decision to classify sites based on the 75th percentile was driven by two competing goals. The percentile must ensure that a majority of visits experience the target performance level, which favors a higher percentile. But it also must not be overly impacted by outliers, which favors a lower percentile. At a high percentile such as the 95th, just a few outlier samples on flaky network connections could skew the classification. At the 75th percentile, a site with 100 visits would need 25 outlier samples to be affected, which is far less likely—while still ensuring that three out of four visits met the target level of performance. Analysis concluded that 75th percentile strikes the right balance.
These thresholds are intentionally broad, designed to apply across the entire web. While they define what is "good", many sites would benefit from optimizing beyond these targets to correlate with their specific business metrics. As the web and user expectations evolve, these criteria and thresholds may be improved and expanded in future years.
Largest Contentful Paint
The LCP thresholds were set with both quality of experience and achievability considerations in mind.
Quality of experience
A commonly cited figure for the limit of a user's focus is 1 second. However, a closer look at the underlying research shows this is an approximation of a broader range. Both Miller and Card, frequently referenced for this topic, describe user tolerance for delays in terms of a spectrum, from roughly 0.3 seconds to 3 seconds. This suggests that a "good" LCP threshold should fall within that range.
Given that the existing First Contentful Paint (FCP) "good" threshold is 1 second and that the Largest Contentful Paint typically occurs after FCP, we further narrow our candidate thresholds to a range from 1 second to 3 seconds.
Achievability
Using data from CrUX, we can measure the percentage of origins that meet our candidate LCP "good" thresholds.
| 1 second | 1.5 seconds | 2 seconds | 2.5 seconds | 3 seconds | |
|---|---|---|---|---|---|
| phone | 3.5% | 13% | 27% | 42% | 55% |
| desktop | 6.9% | 19% | 36% | 51% | 64% |
While less than 10% of origins meet the 1 second threshold, all thresholds from 1.5 to 3 seconds are viable candidates, as they meet our requirement that at least 10% of origins achieve the "good" threshold.
To further verify that a threshold is consistently achievable for well-optimized sites, we analyzed the LCP performance of top-performing sites across the web. We aimed to identify a threshold that is consistently achievable at the 75th percentile for these sites. The 1.5 and 2 second thresholds were not consistent, but the 2.5 second threshold was consistently met.
For the "poor" threshold, we used CrUX data to find a threshold that most origins do not meet.
| 3 seconds | 3.5 seconds | 4 seconds | 4.5 seconds | 5 seconds | |
|---|---|---|---|---|---|
| phone | 45% | 35% | 26% | 20% | 15% |
| desktop | 36% | 26% | 19% | 14% | 10% |
A 4 second threshold would classify roughly 26% of phone origins and 21% of desktop origins as "poor," which falls within our target range of 10-30%. We therefore conclude that 2.5 seconds is a reasonable "good" threshold and 4 seconds is a reasonable "poor" threshold for LCP.
Interaction to Next Paint
The INP thresholds also weigh quality of experience against achievability.
Quality of experience
Research is fairly consistent in showing that visual feedback delays up to about 100 ms are perceived as directly caused by the user's input. This points to 100 ms as an ideal "good" threshold for INP.
Jakob Nielsen's widely cited work defines 0.1 second as the limit for a system to feel instantaneous. This traces back to Michotte's perception of causality research, where participants saw delays of up to roughly 100 ms as a direct cause-and-effect relationship between two moving objects. Delays from 100 ms to 200 ms produced mixed results, and delays over 200 ms broke the perceived connection entirely.
Miller's work is similarly aligned, stating that the delay between depressing a key and visual feedback should be no more than 0.1 to 0.2 seconds. More recent studies on virtual buttons found that delays of 85 ms or less were perceived as simultaneous with the touch 75% of the time. Perceived quality remained high for delays up to 100 ms, dropped noticeably between 100 ms and 150 ms, and reached very low levels at 300 ms.
Given this, we identify 100 ms as the ideal "good" threshold, and 300 ms as a natural candidate for the "poor" threshold.
Achievability
CrUX data shows that the majority of origins already meet a 200 ms INP "good" threshold at the 75th percentile.
| 100 ms | 200 ms | 300 ms | 400 ms | 500 ms | |
|---|---|---|---|---|---|
| phone | 12% | 56% | 76% | 88% | 92% |
| desktop | 83% | 96% | 98% | 99% | 99% |
We paid particular attention to lower-end mobile devices, where they formed a high proportion of a site's traffic, and this further supported a 200 ms threshold. While the quality-of-experience research suggests 100 ms, 200 ms is a reasonable, achievable threshold that reflects real-world conditions.
To set a "poor" threshold, we looked at CrUX data to see which thresholds most origins miss.
| 100 ms | 200 ms | 300 ms | 400 ms | 500 ms | |
|---|---|---|---|---|---|
| phone | 88% | 44% | 24% | 12% | 8% |
| desktop | 17% | 4% | 2% | 1% | 1% |
A "poor" threshold of 300 ms initially looks reasonable. However, contrary to LCP and CLS, INP correlates inversely with popularity — more popular sites tend to be more complex, often leading to higher INP values. Examining the top 10,000 sites reveals a different picture.
| 100 ms | 200 ms | 300 ms | 400 ms | 500 ms | |
|---|---|---|---|---|---|
| phone | 97% | 77% | 55% | 37% | 24% |
| desktop | 48% | 17% | 8% | 4% | 2% |
On mobile, a 300 ms "poor" threshold would label the majority of popular sites as "poor," exceeding our achievability criteria. A 500 ms threshold puts us safely in the 10-30% range of affected sites. The 200 ms "good" threshold is also harder for these popular sites, but as 23% still pass it on mobile, it meets our minimum 10% pass rate.
For these reasons, we conclude 200 ms is a reasonable "good" threshold, and greater than 500 ms is a reasonable "poor" threshold for INP.
Cumulative Layout Shift
The CLS thresholds are based on perception testing and CrUX achievability data.
Quality of experience
Cumulative Layout Shift (CLS) measures how much a page's visible content shifts. Because CLS is a new metric, there was no existing research to directly inform its thresholds. To align with user expectations, we evaluated real-world pages with varying amounts of shift. Our internal testing found that shifts of 0.15 and above were consistently perceived as disruptive, while values of 0.1 or lower were noticeable but not excessively so. Zero shift is ideal, but we identified up to 0.1 as candidate "good" thresholds.
Achievability
From CrUX data, we see that nearly half of all origins have a CLS of 0.05 or lower.
| 0.05 | 0.1 | 0.15 | |
|---|---|---|---|
| phone | 49% | 60% | 69% |
| desktop | 42% | 59% | 69% |
While this suggests 0.05 could be a viable threshold, there are use cases where layout shifts are hard to avoid. For instance, third-party embeds often have unknown heights until they load, causing a shift that can exceed 0.05. The slightly less stringent 0.1 threshold therefore strikes a better balance. It is our hope that the web ecosystem will find solutions for third-party layout shifts, which would allow a stricter 0.05 or 0 threshold in a future iteration of Core Web Vitals.
To determine a "poor" threshold, we used CrUX data to see which threshold most origins do not meet.
| 0.15 | 0.2 | 0.25 | 0.3 | |
|---|---|---|---|---|
| phone | 31% | 25% | 20% | 18% |
| desktop | 31% | 23% | 18% | 16% |
A 0.25 threshold classifies about 20% of phone origins and 18% of desktop origins as "poor." This fits the 10-30% range, so we concluded that 0.25 is an acceptable "poor" threshold for CLS.



