What a decision model does

A decision model produces bounded, structured outputs cheaply and consistently, returning typed answers with probabilities that calling code can act on. A support message can be passed in with a question such as whether it is urgent and which team should own it; the returned probabilities let your code route the ticket, trigger an escalation, or defer to a human. Unlike large language models, which are largely non-deterministic but open-ended enough for reasoning, text generation and tool calls, decision models stay reliable across arbitrary input sets without retraining each time a new classification category appears.

Clef and Clef-flash

Cloudflare is releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI. Clef currently leads the Jev Decision Index, with full results on the live benchmark demo site. Both are fully Jev-API compatible, so swapping in the hosted models is a small change. The weights are also open-sourced on Hugging Face under an Apache 2.0 license for local experimentation.

decision-index-vs-latency.png

How Clef differs from other decision models

Three properties distinguish Clef from the existing field. It carries a vision encoder, so it can classify images as well as text — something Jev does not do today. Its context window is 64k against Jev's 32k, allowing more input state to be classified in one pass. And it scores competitively across the quality evaluations in the Jev Decision Index.

Benchmark

Clef

Clef-flash

Jev

DiffusionGemma Jev

Kev 9B

Laya

BFCL · case exact

98.47

98.76

95.75

96.52

94.51

38.13

ToolRet · nDCG@10

69.19

66.43

65.28

61.21

64.26

12.69

API-Bank · accuracy

91.93

93.11

88.19

83.66

56.30

11.41

Home appliances · case exact

82.95

97.73

52.27

42.05

25.00

0.00

When2Call · accuracy

72.37

65.58

80.97

75.44

49.62

11.94

BANKING77 · macro-F1

94.20

90.93

79.74

74.28

84.83

14.29

CLINC150+OOS · macro-F1

97.43

66.77

89.27

83.49

79.03

3.19

BRIGHT · nDCG@10

45.91

39.26

47.52

42.94

38.53

19.90

Amazon ESCI · macro-F1

57.48

57.39

55.21

53.37

49.22

24.40

PhishNChips · accuracy

79.60

75.05

62.55

85.35

50.75

50.15

On Typesafe's own eval suite, the Clef models beat Jev in three of four areas, with Clef-flash performing notably well for its speed.

Workflow

Clef

Clef-flash

Jev

Invoice processing

64.7

57.1

61.8

Customer service

76.3

77

76.0

Security incidents

62.9

61.7

61.7

Agent trace observability

68.5

69.8

71.6

Latency results across the 43 eval benchmarks Cloudflare ran show the Clef models beating the other decision models on latency, with Laya the exception — very fast, but at a cost in the quality benchmarks above.

Benchmark

Clef

Clef-flash

Jev

DiffusionGemma Jev

Kev-9B

Laya

Median latency  · ms

209.3

38.8

524.1

84.4

51.4

5.8

p95 latency · ms

238.6

122.4

536.0

211.2

187.9

222.5

Hosting on Workers AI adds to that: running on Cloudflare GPUs at the edge keeps network latency low, which makes Clef viable in the hot path for agent decisions, paired with a Workers AI LLM to carry out the resulting action. Clef emits strictly typed outputs like Jev and is fully API-compatible. The larger model is the precision option; Clef-flash targets latency-critical decisions. Cloudflare guarantees that requests and responses are not read, stored or trained on, unless the fine-tuning product is used.

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \\
  -X POST \\
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \\
  -d '{
    "model": "clef",
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {
      "urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
          "billing": "Payments, invoices, and refunds",
          "technical": "Outages, errors, and configuration",
          "sales": "Plans and upgrades"
        }
      },
      "severity": {
        "type": "score",
        "instructions": "How severe is the customer impact?",
        "criteria": ["No impact", "Minor", "Major", "Critical"]
      }
    }
  }'

Training approach

Clef's lineage runs back to an early experiment, published the same week Jev appeared, that adapted DiffusionGemma to emit deterministic probabilities by exposing the logprobs generated by a large language model. That work built on independent research by Matt Mastracci, who has contributed to the ML community and to pull requests against the vLLM inference engine to strengthen DiffusionGemma support.

Clef reuses the concept with a different backbone: Qwen, post-trained for decision use cases. At inference, Clef runs a prefill-only pass on Qwen, then scores the valid schema choices in parallel. The decision step is non-autoregressive, so nothing is generated token by token, which is where much of the speed advantage over autoregressive LLMs comes from. Rather than producing intermediate text to reach a structured answer, Clef and Clef-flash derive schema choices directly from internal backbone representations. That relies on a two-stage attention routing process: each valid choice extracts context relevant to the prompt, letting individual field parameters cross-attend to other fields and back to the original payload before scoring. A lexical prior preserves semantic intent across options. The architecture combines option-specific evidence routing, joint cross-field attention, and schema-bound scoring.

Qwen3.8-27B is frozen for Clef and Qwen3.5-9B for Clef-flash, while the routing head is optimized jointly with rank-256 low-rank adapters. Post-training uses label-smoothed cross-entropy over valid schema outputs together with a Brier loss to refine probability calibration, on internal synthetic datasets that permutate field orders, prompts and schema structures. Reinforcement Learning for Calibrated Decisions (RLCD) serves as a secondary optimization target, granting partial credit to adjacent ordinal choices, rewarding fully precise record outputs, and applying a reference penalty to prevent distribution shift. The result: better classification accuracy, outputs constrained to probabilities rather than generated text, and lower latency than Jev and the base Qwen models.

Fine-tuning Clef

Internal Cloudflare teams have asked for tuned versions of Clef inside agentic workflows — classifying Trust & Safety submissions, triaging Support requests, or deciding inside Bot products whether a crawler is good or bad. Such use cases are narrow, and years of labelled decisions exist to train against. Fine-tuning trades some general-purpose performance for higher accuracy in a specific domain, and Cloudflare's 15-plus years of network data across domains makes that trade attractive: a tuned model can be both more accurate and faster than the generic Clef model. Internal teams are already working on post-training Clef, and these workloads form the remit of the new FDE fine-tuning team and the basis of the RL product.

The RL fine-tuning service

Cloudflare's reinforcement learning product starts as a hands-on offering with its forward-deployed engineer team, which will inform a self-serve platform for capturing data, fine-tuning and redeploying the model on Cloudflare. The plans draw on primitives already in place:

  • Cloudflare AI Gateway — route AI traffic through it to automatically build a dataset of requests for your use case
  • Cloudflare Workers AI — generate rollouts against the base Clef model
  • Cloudflare Containers — RL sandbox for scoring and replaying agent actions
  • [NEW] Trainer — update weights of the fine-tuned Clef model
  • Cloudflare Workers AI + BYO Model — redeploy the fine-tuned model on Workers AI

Several work-in-progress pieces of the AI Platform come together here: AI Gateway capturing AI traffic so customers can use their own request and response data, Containers for RL sandboxes, and Workers AI's Bring Your Own Model (Cog) work, which has progressed since the acquisition of Replicate.

Getting started

The hosted models are documented in the Workers AI developer docs, and the weights are downloadable from the Hugging Face repository. Customers with fine-tuning use cases are invited to apply as design partners.