✉ The Friday AI Brief: the week's 5 best AI stories, tools & comparisons — in your inbox every Friday morning.

Cloudflare logo

Cloudflare Releases Clef and Clef-Flash: Its First Self-Trained AI Models, Built for AI Agents

Cloudflare Releases Clef and Clef-Flash: Its First Self-Trained AI Models, Built for AI Agents

October 2, 2026 — Cloudflare has trained its own AI models for the first time.

Key takeaway: Clef and Clef-flash are open-source “decision models” — they don’t write text, they return structured choices (route this ticket, escalate that one) with probabilities. On Workers AI, Clef-flash answers in under 40 milliseconds, and Cloudflare will let customers fine-tune the models with reinforcement learning.

What happened

On October 1, 2026, Cloudflare released Clef and Clef-flash, the first models trained by its own Workers AI team. They’re hosted on Workers AI, with the weights open-sourced on Hugging Face under the Apache 2.0 license. Alongside them, Cloudflare debuted a reinforcement learning (RL) fine-tuning service so customers can adapt Clef to their own workloads.

What a “decision model” actually is

A decision model belongs to the same family as Typesafe AI’s Jev, the “System One” model that kicked off this category in September. Instead of generating text token by token, it reads an input state and a set of typed questions, then returns a probability for every allowed answer. Your agent gets a structured decision it can act on immediately — route the ticket, block the request, escalate to a human — with no free-form output to parse and no reasoning tokens to wait for.

Clef is the heavyweight: 27 billion parameters, a 64,000-token context window, and support for text, JSON, images, and video (up to four images per request). Clef-flash is the 9-billion-parameter sibling aimed at latency-critical paths. One request can carry up to 64 questions, and both models follow the System One API, so teams already integrated with Jev can switch by changing the endpoint and model name.

The numbers Cloudflare is leaning on

On speed, Clef-flash posted a median latency of 38.8 milliseconds across Cloudflare’s 43 benchmark runs, versus 524.1 ms for Jev — roughly 13x faster. Clef came in at 209.3 ms, about 2.5x faster than Jev. On the BANKING77 classification benchmark, Clef scored a macro-F1 of 94.20, ahead of Jev’s 79.74. Cloudflare claims a Clef model scores highest on 7 of 10 decision benchmarks it ran. Pricing is set at $0.24 per million input tokens for Clef and $0.09 for Clef-flash — output tokens aren’t billed at all.

Why it matters

Here’s the practical angle: AI agents get expensive when every micro-decision requires a full LLM call. Decision models compress the routine stuff — classify, route, gate, escalate — into millisecond-scale, deterministic steps, and they’re cheap enough to leave running on every request. Cloudflare running them on GPUs across its own edge network, close to users, makes the hot path genuinely hot.

The open-sourcing matters more than the benchmarks, though. Jev is a closed API; Clef gives you the weights forever, under Apache 2.0, with an RL fine-tuning platform to make it yours. And Jev-API compatibility means this whole category is starting to look like a commodity layer — one you should be building against, not paying a premium for. For anyone running agents in production, the message of this week is: stop asking your flagship model whether to escalate a support ticket.

Frequently asked questions

Is Clef a chatbot? No. Decision models don’t generate text — they return probabilities for predefined answers, so they’re for the decision steps inside agent workflows, not conversations.

Is Clef really free? The weights are open source under Apache 2.0 and free to download from Hugging Face and run yourself. Cloudflare also hosts them on Workers AI for pay-as-you-go inference.

What’s the RL fine-tuning service? A new reinforcement learning platform that lets customers fine-tune Clef on their own data. Cloudflare is initially working hands-on with design partners before opening it up more broadly.

How is this different from Amazon’s Strands Decider 2B? Similar category, different strategy: AWS went tiny (2B, local-first), while Cloudflare ships two sizes including a 27B model aimed at accuracy, plus hosted inference on its edge network.

Sources: Cloudflare blog, TokenPost, MarkTechPost, Crypto Briefing

Leave a Comment

Your email address will not be published. Required fields are marked *

Get the 5 best AI tools every week

Top AI news, tools, and prompts — one short email. Free, unsubscribe anytime.

Scroll to Top