When the Numbers Jump: Discontinuities in Policy, Markets, and Human Behavior

A sharp threshold in any system—a tax code, a network queue, an admissions policy—creates odd incentives. Around the end of last year, personal finance forums were full of people asking how to deliberately lose money. The goal was to get under the U.S. Affordable Care Act subsidy cutoff at $48,560 for individuals (higher for larger households; $100,400 for a family of four). Details vary by age, location, and plan, but crossing that line could raise health insurance costs by roughly $7,200 per year. Someone earning $55,000 would often be better off reducing their income by $6,440 to qualify for the subsidy than staying above the ceiling.

This is an extreme example, but not a rare kind of case. U.S. tax policy is full of similar discontinuities—the income limits for TANF, Medicaid, and CHIP, for instance—that can make earning more money a net loss. At certain income levels, people are effectively incentivized to shed income, perhaps by buying put options that will expire worthless, transferring wealth to (on average) much richer options traders.

The famous Learned Hand quote defends the individual's right to arrange affairs to minimize taxes. But a system that rewards losing money is sub-optimal. A better design, used for some subsidies, is a gradual phase-out. Slow phase-outs have their own problems, but they avoid the harsh cliff that a sharp threshold creates.

Queues and Network Traffic

A naive queue has a binary behavior: it drops entries when full, and accepts them when not. For bursty network workloads with low overall bandwidth utilization, this can be "unfair," depending on your goals. Random early detection and its variants address this by making the drop probability a smooth function of queue fullness, mitigating the all-or-nothing outcome. The same idea applies, reversed, to link aggregators: there's a sharp discontinuity between the traffic an item gets on the front page versus just off it—effectively a binary "dropped" or "kept" based on a vote threshold.

College Admissions and the Pell Grant Cliff

As universities began using Pell Grant receipt as a proxy for commitment to low-income students, a first-order effect appeared: students just below the threshold had a significantly better chance of admission, while those just above it had a significantly worse one. That seems intended. But look within the groups, and the outcomes invert. Among students who don't qualify, it's the lowest-income who get hurt most. Among those who do qualify, the highest-income get the biggest boost. The second-order effect: savvy parents can game this by shifting income—say, moving money from a Roth to a traditional IRA, or even taking an options loss—to cross the Pell threshold and improve their child's odds at a selective school.

Programmed Outcomes: Elections, Prices, and p-Values

Histograms of Russian election results across polling stations show odd spikes at round numbers like 95% starting around 2004—a pattern consistent with fabricated results that don't bother to look natural. Benford's law is a related tool for flagging unnatural numbers.

In U.S. auto auctions, Mark Ainsworth has pointed out discontinuities at $10k boundaries in both sale prices and the volume of vehicles offered. The effect persists even after adjusting for factors like model year.

Psychology journals show a famously inflated number of papers with p-values just below 0.05. The spike is consistent with several unflattering explanations: authors fudging results, journals preferentially accepting p=0.05 over p=0.055, or authors withholding non-significant findings. Head et al. (2015) surveys this across fields. Andrew Gelman and others have long argued that the entire edifice of statistical significance should be dismantled—not just to remove the cheating incentive but because a bright-line rule for "significance" is a bad way to do science.

Criminal Justice: Drug Amounts on Paper

Histograms of cocaine amounts in federal drug cases show a smooth distribution before the 2010 Fair Sentencing Act, which raised the 10-year mandatory minimum trigger from 50g to 280g. Afterward, a sharp spike appears at exactly 280g. The effect is widely misunderstood: it's not police seizure behavior. The paper's author notes that prosecutors charge amounts that can differ from what's seized, and finds that roughly 30% of prosecutors are responsible for the post-2010 bunching. These "bunching" prosecutors also have distinctive patterns at the 28g five-year threshold, and when one transfers to a new district, other attorneys there start bunching too. The evidence points squarely to prosecutorial discretion, not law enforcement. There is some hint of an anomaly in seizure data (Fig. A8.(c)), but the prosecutorial effect is far larger.

Exams: The 30-Point Fix

Polish language exit exam scores show a curious jump at 30—the passing score—with far fewer scores in the 23–29 range than you'd expect. Math exams show no such effect. An anonymous reddit commenter explains why: graders for the Polish exam, which contains open-ended writing questions, have leeway to "find" a missing point or two to push a failing student to a passing 30. This isn't possible for the rigidly graded math exam, which has no such discontinuity. The bright-line threshold doesn't just create a spike; it changes grading behavior to avoid the severe consequences of failure. A deeper issue is the attempt to discretize a continuous score into a pass/fail certification.

Sports: The Birthday Effect

UEFA Youth League player data shows a strong relationship between birth month and the odds of making a U19 club. Players born earlier in the year—older relative to their age cohort—are heavily overrepresented. The discontinuity isn't in the year-by-year graphs, but if you plot multiple cohorts you'd see a sawtooth pattern: a player born one day before the cutoff can have dramatically different odds than one born one day after. Younger-within-year players who do make it tend to have higher value on the field, consistent with studies showing that discriminated-against groups (e.g., black baseball players post-desegregation, French-Canadian defensemen) need to be better than average to survive selection.

The mechanism is understood: kids are grouped by age-year, older kids are physically more developed, they perform better, and that advantage compounds into higher participation and selection rates. This is arguably a bug in how youth sports work. It persists in part because youth teams aren't feeder teams to pro squads; they have no financial incentive to pick the most skilled-for-age player over the simply older one, making the problem harder to fix than a pro team's locally bad decision.

Collusion in Procurement: Knowing the Floor

Kawai et al. analyzed Japanese government procurement for suspicious bidding patterns similar to those Porter et al. (1993) found in New York. A classic example: a Long Island resurfacing project got bids the DOT deemed too high; the reauction's low bid was 20% higher, and the third auction's was 10% higher still—always from the same firm.

That could just be differing cost structures. So Kawai et al. focused on auctions where the first and second bids were extremely close, making the winner nearly random. The auction rules: a secret bid above a secret reserve wins; if not, the low bid is revealed and a second round is held. In about 97% of those near-random auctions, the first-round lowest bidder also won the second round. The second-lowest bidder stayed second only 26% of the time. Histograms of normalized bid deltas show a sharp asymmetry: the second-lowest bidder almost never lowers their bid by more than the first-lowest did. The pattern persists into third rounds. It's a signature of collusion—bidders learning the floor and refusing to undercut below a certain point.

Restaurant Grades

NYC restaurant inspection scores show sharp discontinuities at 13–14 (the A/B border) and signs of one at 27–28 (B/C). Inspectors have discretion over which violations to tally, and the data suggests restaurants are sometimes nudged up to a higher grade.

Marathon Times: Rounding Up

A histogram of 9,789,093 marathon finishes shows noticeable spikes at every half-hour, plus smaller ones at ":10," ":15," and ":20." Analysis of individual races shows this is partly due to runners speeding up near the end when they're close to a "round" time—a real, if modest, behavioral response to a purely arbitrary threshold.

A Practical Note on Finding and Fixing Discontinuities

These examples aren't just curiosities. Suspicion of discontinuities and figuring out where they come from has real utility, as does applying techniques to smooth them out. The basic toolkit is simple: scatterplots, histograms, CDFs, and time-based visualizations like flamescope. Randomization is a proven way to smooth a sharp threshold in queues, and the same idea can reduce quantization error and other threshold artifacts elsewhere. When you see a sharp edge in the data, it's worth asking why it's there—and who's benefiting from it.