Clef-omni adds audio and video input

Cloudflare has released Clef-omni, a decision model that accepts audio, video, image, and text input in a single call. Where Clef handled text, images, and arrays of video frames, Clef-omni takes audio (wav or mp3) and video (mp4 or webm) alongside those modalities. The model is available through the developer docs and its weights are published open on Hugging Face.

The practical effect is the removal of cascading pipelines: no separate speech-to-text stage, and no need to split audio and image channels out of a video before scoring.

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-omni \
  -X POST \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "clef-omni",
    "state": "Review the installation: a photo of the unit, an audio recording of it running, and a video of the fan.",
    "images": ["data:image/png;base64,<base64-png>"],
    "audio": ["data:audio/mpeg;base64,<base64-mp3>"],
    "videos": ["data:video/mp4;base64,<base64-mp4>"],
    "questions": {
      "label_visible": {"type": "noul", "instructions": "Is the model and serial number label visible in the photo?"},
      "sounds_normal": {"type": "noul", "instructions": "Does the unit sound like it is running smoothly, without rattling or grinding?"},
      "fan_running": {"type": "noul", "instructions": "Is the fan running in the video?"}
    }
  }'

Clef-omni is built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts foundation, which already processed text, imagery, audio, and video in one pipeline. Cloudflare keeps the primary comprehension backbone and discards the text-to-speech output components.

At inference the model runs a single prefill pass over the whole payload, scoring all modalities and candidate parameter options at once. Because Clef models are not LLMs, there is no output token generation step, which removes the overhead of transcribing or captioning media; media elements map directly into the unified sequence, with video and audio synced to visual frames. Candidate values are then pulled from internal embeddings through a two-stage attention routing scheme: each valid option gathers key evidence from the input wherever it appears, and field vectors cross-attend across the full context to produce confidence scores. A built-in lexical grammar preserves option semantics, giving schema-constrained scoring across every input type.

Training follows the Clef recipe: the Qwen3 backbone is frozen, low-rank adapters (LoRA) are trained, and post-training combines label-smoothed cross-entropy loss with Brier score calibration. The stated aim is resilience to schema variations, field ordering, and prompt structures.

Latency stays low per modality — text-only decisions return in about 130 ms at the median, image inputs in about 150 ms, and audio clips in a few hundred milliseconds. A full 21-second video clip with sound is scored in roughly 1.5 seconds, all within one API call. Benchmark results and comparisons against the TypeSafe evals accompany the release.

Benchmark

Clef-omni

Clef

Clef-flash

Jev

BFCL · case exact

98.2

98.47

98.76

95.75

ToolRet · nDCG@10

66.6

69.19

66.43

65.28

API-Bank · accuracy

92.7

91.93

93.11

88.19

Home appliances · case exact

69.3

82.95

97.73

52.27

When2Call · accuracy

63.3

72.37

65.58

80.97

BANKING77 · macro-F1

94.8

94.20

90.93

79.74

CLINC150+OOS · macro-F1

97.7

97.43

66.77

89.27

BRIGHT · nDCG@10

42.0

45.91

39.26

47.52

Amazon ESCI · macro-F1

57.8

57.48

57.39

55.21

PhishNChips · accuracy

73.2

79.60

75.05

62.55

Workflow

Metric

Clef-Omni

Clef

Clef-flash

Jev

Invoice processing

Exact actions

60.2

64.7

57.1

61.8

Invoice processing

Primary action

82.0

86.2

73.3

83.1

Customer service

Exact actions

71.6

76.3

77.0

76.0

Security incidents

Exact actions

61.7

62.9

61.7

61.7

Agent trace observability

Primary action

65.8

68.5

69.8

71.6

Pricing changes across the family

Clef-flash is now cheaper than Jev, following optimizations to the model. Pricing below covers input tokens; the developer docs carry current figures and explain how image and audio modalities convert to input tokens.

  • Clef-flash — previously $0.09 per M input tokens, now $0.038 per M input tokens
  • Clef — unchanged at $0.24 per M input tokens
  • Clef-omni — launched at $0.15 per M input tokens

The lower price comes with a trade-off: the hosted Clef-flash context window drops from 64k to 24k. The Hugging Face weights are untouched and were trained for a 256k context window for self-hosting. Cloudflare reports that only 0.24% of requests exceed 24k input tokens, which motivated the reduction; Clef keeps its 64k window, and users needing more context are directed to it.

Serving optimizations for Clef

Clef itself got faster through changes at the serving infrastructure layer rather than new weights or architecture changes, so no new model weights were released. Clef now runs on SGLang; Cloudflare worked with the SGLang team and contributed Clef support in PR #42721, due in SGLang 0.5.22.

Input size

Before: median / p95 (ms)

Now: median / p95 (ms)

Median speedup

~800 tokens

262 / 438

152 / 351

1.7×

~3,400 tokens

616 / 777

305 / 531

2.0×

~16,000 tokens

2,721 / 3,250

1,635 / 1,805

1.7×

Self-hosting Clef remains possible with the released weights, and updated SGLang launch commands are available in the Hugging Face repo and the SGLang cookbooks. Clef is Jev-API compatible and reachable via AI Gateway, so switching is a matter of changing the model ID.

Deployments in production

Inside Cloudflare, the appeal of Clef is that detection and classification move into the model layer. Teams previously might have needed a specialized ML team and a domain-specific corpus, since even zero-shot classifiers were often not powerful enough to generalize without tuning; the model can now be called directly on input data.

Reported uses include the public GitHub docs repo detecting and closing spam issues, the EmDash content management system moderating plugin libraries for phishing, the data loss prevention team scanning for PII such as government IDs, and the threat intelligence team detecting malicious domains.