Tag: Evals

Jev-Like Models

TypeSafe’s Jev has sparked a lot of community activity. I tried to collect a few and list them here:

The blossom of these Jev-like models shows there are huge technical challenges in matching Jev’s bench results, though Jev still leads on accuracy and speed. But Jev is not irreplaceable. For example, you can trade accuracy for speed with Reflex.

The bench has a full list of more Jev-like models you can check out.

Open Takes on Jev: SemIf and Laya

OpenJev has been renamed SemIf.

It’s wise to separate it from Jev’s marketing. It’s an independent research project that just provides the same API. The underlying architecture might be completely different, since TypeSafe never disclosed Jev’s own model or training.

SemIf uses Qwen and MiniCPM5 as its core models. Instead of an autoregressive decoder, it does one forward pass and reads probabilities over filtered logits.

Another project, Laya, is also worth a look. It claims to outperform Jev by a tiny margin. The model is small, 421M params. Laya evaluates typed questions (choice, score, noul) over any state, text, email, ticket, or a JSON document, in a single forward pass too.

I ran it myself on an M4 Pro. Without preloading, it took 66 seconds to classify a code-change type: given a PR diff, check whether it’s a feature change, a docs change, a test change, an IaaS change, a version bump, and so on. I gave it a version-bump-only diff, and it scored version bump at 0.7477, with everything else below 0.55. It classifies well. It already looks usable.

Testing Jev: Three Playground Runs

I got early access to Jev after writing about it 2d ago. Here are three tests I ran in the playground, with the actual state, questions, and answers.

Is a Jedi sandwich a sandwich?

State:

{
  "food": "Jedi sandwich",
  "definition": "Put Luke Skywalker in between two slices of toast."
}

Question:

{
  "is_sandwich": {
    "type": "noul",
    "instructions": "Is `food` a sandwich?",
    "criteria": {
      "true": "A sandwich is a food dish where a filling, such as meat, cheese, vegetables, or spread, is placed between structural starch",
      "false": "The food has no bread enclosing a filling or uses only a single slice of bread, or uses a non-bread wrapper such as a tortilla, wafer, or cookie."
    }
  }
}

Answer:

[…723 words]

How Do You Prove a Coding Skill Works? A Ponytail Case Study

OpenAI wrote a post called Testing Agent Skills Systematically with Evals. The point is simple: if you only judge a skill by feel, you won’t notice when it gets worse. You need a test set, a pass/fail check, and a score you can repeat.

Ponytail is a good example to study. It’s a skill that makes an AI agent write less code. Its claim is easy to test, so its benchmarks folder is a good case study of the same idea, but for a coding skill instead of an agent skill.

[…849 words]

AI Foundations 6 - Prompting and Evaluation

The exact same model can look brilliant or mediocre depending entirely on how you ask it something. Swap a vague instruction for a specific one, and the quality, the format, even the correctness of the answer can shift noticeably — without touching the model at all.

Prompt structure

A prompt that works tends to spell out the same handful of things: what the task actually is, what context matters, and what shape the answer should take. Keeping the role, the task, any constraints, and any examples visibly separate makes a long prompt easier for the model to parse — and easier for you to debug later.

[…317 words]

Early Reactions to GPT-6 Astra

It’s been about two days since GPT-6 Astra launched, and enough people have access now that useful early reports are starting to show up.

The strongest feedback I’m seeing is that Astra improved more on agency than intelligence.

Artificial Analysis currently gives Astra the same Intelligence Index score as GPT-5.6 Sol: 61. Its Coding Agent Index moves from 65 for Sol to 67 for Astra. That’s an improvement, but hardly the generational jump the GPT-6 name suggests. (Hacker News)

[…305 words]

AI Engineering SKill Map

Andrew Ng’s wrote an AI Engineering Skills Map, which lists six things to learn: LLM foundations, grounding models with data, building agentic systems, evaluation-driven development, operating in production, and machine learning foundations.

As he broke it down, I think the most important skill is learning how to build reliable systems from LLM’s uncertain behavior.

You don’t know in advance what an LLM will output

This is the only only truth you need to take away in this post, if you can’t remember all.

You can’t design AI software the way you design normal software — plan it, build it, ship it — because you can’t plan around an output you haven’t seen yet.

The AI engineering stack is packed with jargons now: MCP, CLI tools, sandboxes, memory and context management, harness, loop, RAG, prompt, multi-agent, (sorry, I cannot name them all, too much). But the core of AI engineering is surprisingly simple:

Build something, observe what it does, evaluate whether that’s good enough, change the weakest part, and repeat.