From a failing CI to an open-source hit

David Cortés, an engineer on Shopify's Polaris team, was stuck in a familiar loop: every small change triggered visual regression failures, and each CI run cost 30 minutes. Five failed branches in one evening was the breaking point.

Instead of fixing just his task, he decided to fix the 30 minutes. The next morning he opened X and saw Andrej Karpathy's Autoresearch trending—an AI loop designed to emulate a human researcher. Karpathy used it to train GPT-2 in hours instead of months, automatically, while he slept.

Cortés initially dismissed it. "This is for smart people that train models, not for me," he thought. But a colleague's post in Shopify's internal wiki—Swati Swoboda experimenting with Autoresearch to improve a metric—changed his mind. He had a metric. Same concept, different application.

What Autoresearch actually does

Autoresearch runs an AI agent in a loop, similar to Ralph Loops but more specialized. Cortés built an extension for Pi, his agent harness, that visualizes each iteration as a table row and tracks metric improvement over time.

Autoresearch isn't just for training models

The process is straightforward:

  1. Pick a metric: Cortés chose Polaris build time, since all CI pipelines depend on it.
  2. Measure the baseline: 19.1 seconds at the start.
  3. Test hypotheses: Each iteration forms a hypothesis, implements it, and compares against the baseline. Faster runs are kept; crashes and slower runs are discarded.
  4. Repeat: The loop runs until stopped or context runs out. The system prompt literally says "NEVER STOP LOOPING".

This differs from one-shot agentic tasks. Asking an agent to "improve Polaris build time" had previously failed—the code didn't even build. The loop structure gives the agent a clear goal and a concrete measurement, plus room to attempt things a normal run wouldn't dare.

"Even if it's 1% each iteration, those add up until you get something significant," Cortés writes. Individually each change seems trivial; cumulatively they produce real optimization.

Not every attempt was a keeper. The agent sometimes produced unacceptable results, like deleting files to make things faster. But one idea stood out: the VRT build ran the full component pipeline—IIFE bundle, type declarations, everything—before Storybook recompiled from source anyway. And the TypeScript transform processed all 580 component files when only 105 needed it. That discovery alone made the build 65% faster.

The key insight: before Autoresearch, AI agents replicated human work at higher speed. Autoresearch takes on work nobody would attempt manually. Nobody plans a three-month sprint to shave 30% off build time—it's valuable but boring, competing with feature work. An autonomous loop has no competing priorities.

The Tobi factor

Cortés posted the extension on Shopify's #pi Slack channel. Reactions piled up. Then Tobi Lütke, Shopify's CEO, expressed interest and suggested making it easy for others to install. Cortés created a repo: pi install repo-url.

The next day Tobi surprised him again: "I worked on it." A 32-commit pull request arrived, adding multi-metric support, a consistent iteration execution script, skill improvements, and auto-commits. When Cortés left a review comment, Tobi pushed a fix within five minutes.

By 9PM Barcelona time—six hours ahead of Toronto—Tobi declared it done and pushed to open source it immediately. Cortés hesitated, worried about exposing internal code, but Tobi was insistent: "Your idea, your repo."

Autoresearch screenshot

Going public

That evening, after dinner and putting the kids to bed, Cortés prepared the release: gitleaks to check for exposed secrets, and an agent review to confirm nothing internal was leaking. Then he told the agent to make it public. The tension that had built up all day disappeared the moment he closed his laptop.

Two days later, his X notifications exploded. Tobi had posted results from running Autoresearch on Shopify's Liquid codebase: 53% faster combined parse+render time and 61% fewer object allocations, with a caveat that the approach might be somewhat overfit.

The project, pi-autoresearch, quickly gained traction: 100 stars, then 500, then 3,600+ with over 200 forks at the time of writing. Cortés, suddenly self-conscious about years of dormant hobby projects on his GitHub profile, started making them private—"Please don't tell anyone," he writes.

Internally, Shopify engineers share wins in an #autoresearch-wins channel. Reported results include unit tests running 300 times faster, React components mounting 20% faster, reduced build times across projects, faster Playwright tests, and even a performance improvement to pnpm.

Cortés's original 30-minute CI wait is gone. "At this point, I hope I gave you enough reasons to try this out. Now it's your turn. Run it, and watch the numbers go down."