Vercel's AI SDK for Python now ships an experimental evaluate() API that talks to Jev, a model built around multiple-choice answers rather than free-form text. You supply data and a set of questions; Jev returns its selections along with a confidence value for each. It makes mistakes like any model, but it makes them quickly, and its intended scope is narrow decisions — whether it is accurate enough remains a question each team has to test for itself.

What Jev actually is

The closest familiar analogy is a universal classifier. Classical classifiers require labeled data per task: a spam filter learns from emails tagged spam or not spam, and every new domain means a new dataset and a new round of training. Jev sidesteps that by starting from a large language model, which already carries broad human knowledge in its weights, then behaving like a classifier rather than a text generator.

The practical consequences are that it is cheap and fast, it emits structured JSON, and its answers are constrained to the types and choices declared in the questions you pose.

The evaluate() API

Jev's official API is deliberately small, and the Python SDK reproduces it almost verbatim. There is one function, evaluate(), plus a few supporting types. It takes a model, a state (a plain string or arbitrary JSON), and a mapping of questions.

A question is one of three types:

  • ChoiceQuestion — pick one answer.
  • ScoreQuestion — rate the state on a scale you define.
  • NoulQuestion — estimate the probability that a statement holds.

Installation is uv add ai. You also need an AI Gateway key exported as AI_GATEWAY_API_KEY. Full signatures and types live in the reference docs, which the SDK authors suggest pointing your coding agent at.

import asyncio

import ai

from ai.ops import experimental as jev

async def main():

result = await jev.evaluate(

ai.get_model("typesafe-ai/jev"),

"You're Neil deGrasse Tyson",

{"bigger": jev.ChoiceQuestion(

instructions="Which is bigger, the Sun or the Earth?",

criteria={"Sun": None, "Earth": None},

)},

)

print(result.value["bigger"])

if __name__ == "__main__":

asyncio.run(main())

Case study: distinguishing Python from English

The motivating use case is a classifier proper. The author previously wanted an agentic Python REPL that could tell, without an explicit mode switch, whether the line being typed was English or code. Building that classifier by hand — training on a random sample of Python and English — took two days and produced jittery behavior: if i was highlighted as Python, if i i flipped to English, and if i is flipped back. A problem that looks trivial turned out to be hard.

Jev, asked the same question — does this look like English or Python to you? — fares noticeably better than that hand-rolled attempt, but the gaps persist. The string "what's" + " up reads as English to Jev, even though it is plainly a half-typed Python expression. The surrounding code is essentially the opening snippet with a different question.

Case study: making Jev write Python

Jev cannot generate prose, so getting code out of it means reducing generation to a sequence of choices. One approach picks the next character from a set of letters, punctuation and whitespace; another selects one word at a time from a vocabulary of some N English words.

An initial attempt with "letters + Python keywords + whitespace" could barely produce an if statement. The workable alternative was to have Jev construct an abstract syntax tree instead:

  1. An LLM expands the user's prompt into a concrete plan for Jev. Without a detailed plan, Jev struggles with even elementary tasks.
  2. Jev builds the AST through successive choices. Each step shows the current program, marks the field being filled, and offers candidate nodes — a function call, an arithmetic operation, a variable — each with a preview of the resulting code.
  3. The host applies each choice to the tree and renders it back to Python, handling punctuation and indentation itself while Jev decides structure and content. Repeat until the tree is complete.

This produced syntactically valid but mostly incorrect Python. The conclusion is blunt: classifiers are a poor substitute for an LLM on generation tasks. The example implementation is published as jev_python_ast.py for anyone who wants to improve on it.

Getting started

Both experiments fell short of their goals, but Jev remains worth exploring as a new class of model. The entry points are the same: uv add ai and an AI Gateway key. A copy-paste prompt for a coding agent:

Create a minimal playground for Jev using Vercel's AI SDK for Python. Initialize the project with uv (Python 3.12 or later) and install the SDK with uv add ai. Use ai.get_model("typesafe-ai/jev"). Write main.py that reads input line-by-line in a loop, calls ai.ops.experimental.evaluate() with a ChoiceQuestion to classify each line as Python or English, and prints the choice, probabilities, confidence, and latency. Exit cleanly on Ctrl+C. SDK docs: https://ai-python.dev/docs/reference/ops#experimentalevaluate Prompt the user to create an AI Gateway key and explain how to set the AI_GATEWAY_API_KEY env variable. After that, the user can run uv run python main.py. Gateway key setup docs: https://vercel.com/docs/ai-gateway/authentication-and-byok#quick-start