A GPU-Powered Path Between Hugging Face and Cloudflare
Cloudflare and Hugging Face are joining forces to reduce the friction developers face when deploying AI models. The partnership, announced today, will weave Hugging Face’s catalog of more than 500,000 models—which serve over one million downloads daily—directly into Cloudflare’s developer platform. The two companies plan to roll out three major capabilities over the next few months.
- Serverless GPU models on Hugging Face: Developers will be able to run models without provisioning or paying for idle infrastructure.
- Optimized models in Cloudflare’s catalog: Popular Hugging Face models will be available natively within Cloudflare’s model lineup.
- Cloudflare integrations in Hugging Face Inference: Cloudflare will become part of Hugging Face’s suite of inference solutions.
“Hugging Face and Cloudflare both share a deep focus on making the latest AI innovations as accessible and affordable as possible for developers,” said Clem Delangue, CEO of Hugging Face. “We’re excited to offer serverless GPU services in partnership with Cloudflare to help developers scale their AI apps from zero to global, with no need to wrangle infrastructure or predict the future needs of your application—just pick your model and deploy.”
Deployment From Either Side of the Fence
The integration is designed to meet developers where they already work, regardless of which platform they start from. Cloudflare builders will eventually be able to select Hugging Face models—optimized for performance and speed—directly from the Cloudflare dashboard and deploy them as part of their applications.

Conversely, developers who prefer browsing the Hugging Face Hub will be able to deploy their chosen model directly from the Hugging Face interface to Cloudflare’s Workers AI runtime, bypassing the need to manually export configurations or set up separate infrastructure.

Expanding Inference Options on Hugging Face
Hugging Face currently offers several ways to serve predictions without managing servers, ranging from the rate-limited, free Inference API to dedicated hardware via Inference Endpoints, as well as in-browser execution through Transformers.js. Cloudflare’s involvement will extend these options to include new serverless GPU inference experiences powered by Cloudflare’s network. Specific details and availability are still under wraps, but both teams signal that this is only the first step toward deeper edge-computing use cases.



