Skip to main content

Tech blog

What we can learn from forcing Jev to code

What is Jev?

Jev was launched on the 15th of September 2026, and rapidly became the AI community's latest fascination. Why? Unlike the current wave of LLMs, decision models like Jev don't produce text (or other media) output, they only answer questions. In return, however, answers are very fast, very cheap, and are guaranteed to match an expected structure. This opens up a world of new applications to which generative LLMs are not well suited.
TypeSafe, the makers of Jev, go to quite some lengths to explain that Jev doesn't produce free-form outputs, and shouldn't be used for generative use cases... but what if we did it anyway?

How can we make a decision model write code?

Jev gives us three primitives, all of which work with a 'state' (i.e. the input):
  • Noul
    : Evaluate a question and give the probabilities that the answer is yes or no
  • Choice
    : Evaluate a question and give the probability distribution across a fixed set of options
  • Score
    : Evaluate a rubric and give a position on a spectrum
On the face of it, those primitives don't give us a way to ask Jev to write code for us, so we'll have to get a little creative.
The first thing to note is that code follows a strict grammar: it follows a well-defined structure, and at any given point in a program, there is only a fixed set of options for what might come next. Given that fixed set of options, we can use a Choice question (with a description of the program we want to produce as state) to ask Jev for the next part of the program. If we repeat that, over and over, building up the program as we go, eventually Jev will write us some code!
There are a few different ways to approach this: we could ask for the next character and ask Jev to type it out one "keystroke" at a time, or we could ask for the next syntax token (`for`, const, return), or we could ask for the next node in an abstract syntax tree (aka an AST). That last option is pretty similar to the second, but it allows us to choose concepts like 'a while loop' in one go, and then come back to fill in the blanks (the condition, the body of the loop, etc) later.
For my experiment, that's the route I chose - but it would be interesting to go back and try the other approaches too, and see how they compare! I also chose to use a subset of TypeScript, but the principles here apply to any language.
To actually write a program, we can use a beam search. The generation loop looks roughly like:
  1. For every live branch, look at the sequence of decisions that define that branch and use it to build an AST and print that as text.
  2. Put the user-provided spec and the work in progress AST into state. Identify the spot in the AST we're looking at filling in with a marker.
  3. Generate all the possible next AST nodes that could legally go in that spot. This is good old-fashioned deterministic code!
  4. Ask Jev a Choice with those nodes as options.
  5. For each option, add it to the decisions so far to build a new branch to explore. At the same time, record the probability Jev assigned that option. We use the mean of the log of the probabilities of the decisions to score a branch - i.e. as a heuristic of how confident Jev was in the choices that make up this branch. Add the new branch to the next round's candidates.
  6. Once every live branch has been extended, keep only the top K (as ranked by the above heuristic).
This allows us to find a balance between a greedy algorithm (always picking Jev's highest rated choice) and an exhaustive algorithm (exploring all choices at each step, up to some limit since programs can be infinite).
At the branch nodes of the AST, we're generally asking Jev to pick syntactic structures (a for loop, an if block, etc), which is fairly straightforward. At the leaves of the AST, though, we have three types of node to ask Jev about, depending on the context:
  1. Symbol references (e.g. an existing variable), where we can provide Jev with all the valid options.
  2. Literals (e.g. a number or a string). In our simple experiment we constrained this to literals we could extract (via regex) from the user input, plus a handful of simple literals that often come in handy (0, 1, true, false, etc).
  3. New names (e.g. a new variable). Here we have to ask Jev to spell something out one character at a time, providing the valid characters as choice options (plus a special STOP option).
At the end of all this, we end up with a series of candidate programs, and we need to decide which one to pick. To do so, we can turn to Jev again. We run each of the programs, record the output, and ask Jev a set of Noul questions: given the user's input spec and the program's output, does the program do what the user wanted? Jev gives yes/no answers with associated probabilities for each of those questions. We throw away all the candidates with a 'yes' probability under a threshold, and if we have any candidates left we pick the one with the highest 'yes' probability; if all of the candidates fall below the threshold, generation has failed.
You may notice that this is similar to how generative language models work, with one difference: here, we wrote the loop ourselves around a model that wasn't trained for it. A base LM is also a classifier: it predicts the next token, and only becomes a generator through a loop that picks a token, appends it and asks again. Restricting that loop to legal choices is also an established practice (opens in new tab), as is building code as a grammar-driven syntax tree (opens in new tab). So what we're really testing is the model inside the loop. Jev wasn't designed for generation, so how does it cope?

Does it work?

Yes and no. It feels somewhat like sorcery, but this approach does produce TypeScript programs which parse and execute! For simple prompts, those programs even give the right answer. For more challenging prompts (e.g. that need a modicum of planning), however, the approach tends to fail. During my experimentation, I never managed to generate a correct program for "print the first 5 Fibonacci numbers" - so frontier coding models needn't worry just yet.
We should be fair to Jev, though: TypeSafe did tell us not to do this, so it's no surprise that results are limited.
You can see the code I used for this experiment here (opens in new tab).

So what is a decision model good for?

Decision models might not be great at coding, but they're still an interesting and valuable addition to an AI developer's toolkit. They excel at making rapid judgements at very little cost - all of my experimentation cost well under a dollar. Use cases like guardrails around LLM output, classification of unstructured data, and even real-time loops (e.g. driving browser interactions) are all very much worth considering. Even in this experiment here, we can see Jev doing a good job at judging whether a program's output meets a spec's requirements (correctly judging the outputs from all 10 of the evaluation cases I tried).
During the course of trying this out, I came across a few things I thought were interesting:
  • The decisions are actually only a small part of the harness code and the overall algorithm. Most of the complexity remains in the deterministic code (computing valid next tokens, building and printing the ASTs, etc). This should probably be expected: decision models like Jev work well as a sprinkle of magic used to guide more traditional code, rather than doing the bulk of the work.
  • When asking Jev a Choice, the answer probabilities sum to 1. My initial thought was to use a fixed threshold to cut off unpromising branches, but in situations where there are many plausible next nodes (e.g. the very first node in the AST), each can receive a relatively low score, which a fixed threshold may reject.
  • A chain of independent decisions is not the same as a plan: even when Jev makes a series of confident choices about which AST node to pick, the result can be useless. LLMs, especially reasoning models, can plan ahead and maintain that context much more effectively.
  • A solution is only as good as its algorithm. Jev was pretty effective at answering a Noul question about whether a partial program was "on track", so my expectation was that building this into the beam search heuristic would improve the quality of output. When I tested that approach, however, I found that the harness tended to churn through spurious statements (e.g. console.log calls doing nothing useful) until it found something that Jev happened to score well. This underscores the need to test your hypotheses, as ever.
Similarly, I looked into using a large set of Noul questions (in a single request), one for each AST node option, rather than a Choice question. This avoids the above problem of multiple plausible candidates reducing each other's probabilities, but the result was programs padded with plausible but ineffective nonsense (e.g. if (false && (true || false)) without an increase in actual effectiveness.
Overall, I really enjoyed pushing Jev to its limits, and I'm excited to use decision models like Jev more in work in future - but I'll be sticking with more traditional models for coding, for now!

Rowan Hill

Senior Technical Principal

07 October 2026

Subscribe
to our monthly newsletter for our latest expert content.

About the AuthorRowan Hill

Rowan is one Softwire’s most experienced technical leaders, having been with the company since 2007. With a career spanning technical, delivery, operations, and recruitment roles, he brings a uniquely broad perspective to every project. As a Technical Principal, he recently played a key role in the leadership team of a high street bank’s first SAFe value stream - initially acting as both Technical Lead and Domain Architect before transitioning to a supporting role as permanent hires were made. During the bank’s digital transformation, he worked closely with the Technical Design Authority to improve documentation standards and design review processes. Beyond client work, Rowan led Softwire’s summer internship programme for several years. Under his guidance, intern teams developed software for charities and socially responsible causes - making a real-world impact while bringing fresh talent into the industry. He also keeps an acoustic guitar in the office, though he insists he’s not really that good.