
A single Python function as a Jev-like wrapper for LLMs, vision models included
A blog post from September 2026 shows a single Python function that turns several LLM providers, vision models included, into one uniform call interface and returns probability-scored answers to questions written in plain text.
A blog post published in September 2026 describes a single Python function that turns several LLM providers, vision models included, into one uniform call interface. The author says the idea grew out of reading about Jev and the self-hostable projects around it, OpenJev and SemIf, which introduced him to the trick of reading an LLM's token probabilities.
What "Jev-like" means here
Jev's documented request format carries a JSON document with a text state and a set of questions. The author's wrapper keeps that structure but adds a field of his own: an attachments entry for images, which the documented format does not cover. So "Jev-like" refers to the request and answer structure, not to the project itself.
How the wrapper works
Every question is rewritten as a multiple-choice prompt whose options are lettered from A to T, then sent as an ordinary HTTP request. The model is forced to emit a single token, and in exchange the API returns the letters together with their log probabilities. Those numbers are converted into probabilities, and the answer is derived from them: a choice question goes to the option with the highest probability, a noul (yes/no) question returns the probability of true, and a score question returns the expected level between its criteria. Each question accepts 2 to 20 options; if a provider leaves out an option that is not negligible, the function raises an error instead of guessing.
Providers, images and speed
The example talks to a local llama.cpp server by default and to OpenAI when a key is present; the API key is sent only to api.openai.com. llama.cpp needs Chat Completions for log probabilities, while OpenAI needs the Responses endpoint to expose enough alternatives, and the script handles that difference itself. Images travel as base64 data URLs, so the same questions can be asked about a webcam frame. In the demo the author reads frames with OpenCV, which is there only to reach the camera, and prints a table of answers. Gemma 4 12B on an RTX 3090 gives about one frame per second with three questions per frame; gpt-6-luna over the API gives about 0.2 frames per second.
What the author warns about
The author is explicit that specialized computer vision models are far more efficient, and that the appeal of his approach is flexibility: a new condition is added by describing it in plain text. He also notes that processing the input still costs time, that a shared state prefix can be cached if the backend supports it, and that his throughput figures come from an unoptimized setup with a separate connection per question per frame.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.