Article

Sep 28, 2026

What Is Jev? TypeSafe AI's System One Model, Explained for PE Operators

A new model by TypeSafe called Jev was recently released and gained traction in the tech community. This article outlines what Jev is and how to implement it in current pipelines or process.

TLDR

Jev is not a chatbot. It is a new kind of AI model from TypeSafe AI that answers narrow questions about your data with a typed value and a confidence score, not a paragraph of text.

That makes it a strong fit for the high-volume judgment calls buried inside portfolio company workflows: routing support tickets, coding invoices, scoring leads, triaging documents. For operating partners and portfolio company leaders, the value creation case is simple: AI automations that are faster to build, cheaper to run at volume, and easier to maintain.

What is Jev?

Jev is the first public model from TypeSafe AI, released in limited early access on September 15, 2026 alongside a $40 million seed round led by DCVC. TypeSafe is a San Francisco lab founded in 2024 by Diogo Almeida, Erik Gafni and Sasha Sheng. Almeida spent roughly four years at OpenAI working on RLHF, InstructGPT, ChatGPT and GPT-4.

In the two weeks since launch, Jev has drawn real attention across the developer community. TypeSafe calls it the first "System One model." That label is the easiest way into what Jev does, so it is worth a short detour.

System One vs. System Two thinking

The terms come from psychologist Daniel Kahneman, who popularized them in his 2011 book Thinking, Fast and Slow. My summary is brief; the book is worth reading in full.

Kahneman describes two modes of thought. System One is fast and automatic: reflexes, routines, gut reactions. If someone jumps out and startles you, you don't stop to ask who they are or where the nearest exit is. You flinch, shout or push back before you've consciously thought at all.

System Two is slow and deliberate: solving a puzzle, working through a complex negotiation, building a financial model. It takes effort, and it's where reasoning happens.

The AI community has borrowed this framing. Frontier LLMs such as GPT, Claude and Gemini are increasingly positioned as System Two tools, built for reasoning, writing and multi-step problem solving. Jev is positioned as the opposite: a System One model built for fast, well-scoped judgments that a knowledgeable person could make in a few seconds.

How Jev differs from an LLM

An LLM generates text, one token at a time, in whatever form it likes. Jev does not generate text at all. Functionally, it behaves like a classifier: you give it context and a well-defined question, and it places the answer in a bucket you defined.

Classifiers are not new. Statistical and Bayesian classifiers have been in use for decades, and you interact with one every day: your spam filter. Each incoming email is scanned for signals, and based on what's present, it lands in your inbox or your spam folder. What's new with Jev is pairing that decision-first design with the language understanding of a modern foundation model.

You send Jev a state (the content or context) and one or more questions. Each question uses one of three types, and each returns a fixed, predictable shape:

Question type

What it does

What it returns

Choice

Picks one option from a list you define

choice, probabilities, confidence

Score

Rates the state on a rubric you define

score, probabilities, confidence

Noul

Answers a yes/no statement

noul (a probability from 0 to 1)

Source: TypeSafe documentation

You can mix all three types in a single call. TypeSafe evaluates each question independently and in parallel, so adding questions barely changes response time.

The way I describe it: Jev is not an LLM, but an LLM can be Jev. With enough prompting and guardrails, an LLM can do everything above. Jev simply does it natively.

Why it matters: the hidden cost of making LLMs behave like software

The biggest cost in most production AI automations isn't the model. It's the code around the model that forces its output into a shape the rest of the system can use.

Since OpenAI opened its ChatGPT API in 2023, developers embedding LLMs into applications have found themselves writing the same instructions over and over: return valid JSON, return only a whole number, don't add commentary. Often they're asking more than once in the same prompt.

Say you want a model to score inbound customer emails from 1 to 5 and feed that score to a dashboard. If the model returns "4" as text instead of 4 as a number, or adds "Score: 4 out of 5," the pipeline throws a type error. Programming languages enforce this through type safety (See what Jev did there?).

Teams handle this with parsing, validation and retry layers. Done diligently, those layers work the overwhelming majority of the time. But they consume a disproportionate share of development time, and they need maintenance: a new model version can break logic that used to work. Most major LLM providers now offer native structured-output modes that close much of the formatting gap, so the real question is what else you get.


Frontier LLM

Jev

Output

Free-form text, optionally schema-constrained

Typed answer from options you define

Can return an off-list answer

Yes, without constraints

No, by design

Confidence signal

Not native

Probability per option plus confidence

Speed (vendor-reported)

Seconds for longer outputs

About 100 ms for most calls

Input pricing (vendor-reported)

Varies by provider

$42 per billion tokens; output free

Best for

Writing, summarizing, reasoning

Classifying, scoring, routing, yes/no checks

Sources: TypeSafe, MindStudio, Flavio Copes

For a portfolio company, that translates into three value creation levers:

  • Lower build cost. Less code to guard outputs means engineering hours go to the workflow, not the plumbing.

  • Lower run cost at volume. Classification tasks that run thousands of times a day become cheap enough to automate end to end.

  • Built-in control points. A confidence score gives you a natural threshold for when to act automatically and when to escalate to a person, which is exactly what finance, compliance and audit teams want to see.

One caution on benchmarks: when Jev is compared against LLMs, check whether the LLM baseline was given a structured-output layer. That changes the comparison.

Where Jev fits in a portfolio company

Jev falls wherever a workflow repeats the same narrow judgment hundreds or thousands of times. Most middle-market portfolio companies have several of these hiding in plain sight.

Function

Workflow

Jev question type

What it replaces

Finance and AP

Code incoming invoices to GL accounts

Choice, gated by confidence

Manual coding or brittle rules

Customer support

Route tickets and flag urgency

Choice + Noul

Tier-1 triage by hand

Sales

Score inbound leads against the ICP

Score

Rep gut feel or static lead scores

Customer success

Score sentiment on customer emails for a churn dashboard

Score

Surveys and anecdote

Compliance

Flag documents containing sensitive data

Noul

Keyword searches

Deal team

First-pass screening of teasers and CIMs

Several Scores, weighted in code

Associate hours on low-fit deals

Writing, summarizing and multi-step reasoning still belong to System Two models. The strongest pattern is to use them together: Jev handles the fast decisions (what is this, where does it go, is it safe) and hands only the cases that need depth to an LLM or a person.

How to implement Jev: Python examples

Calling Jev looks a lot like calling any other model API. I've used Python because it's the most readable of the widely used languages. You'll need Python 3.10 or later and an API key from the TypeSafe console, set as the TYPESAFE_API_KEY environment variable.

Example 1: Auto-code invoices, escalate when unsure

from typesafe_sdk import Choice, TypeSafeClient

client = TypeSafeClient()

GL_ACCOUNTS = {
    "5100": "Cost of goods: raw materials and inventory",
    "6100": "Software and SaaS subscriptions",
    "6200": "Professional services: legal, accounting, consulting",
    "6300": "Travel and entertainment",
    "6400": "Facilities: rent, utilities, maintenance",
}

AUTO_POST_THRESHOLD = 0.85  # tune against a sample your controller has already coded


def code_invoice(invoice_text: str) -> dict:
    response = client.system_one(
        state=invoice_text,
        questions={
            "gl_account": Choice(
                instructions="Which general ledger account should this invoice be coded to",
                criteria=GL_ACCOUNTS,
            ),
        },
    )
    answer = response.answers["gl_account"]
    route = "auto_post" if answer.confidence >= AUTO_POST_THRESHOLD else "ap_review_queue"
    return {
        "gl_account": answer.choice,
        "confidence": answer.confidence,
        "route": route,
    }
from typesafe_sdk import Choice, TypeSafeClient

client = TypeSafeClient()

GL_ACCOUNTS = {
    "5100": "Cost of goods: raw materials and inventory",
    "6100": "Software and SaaS subscriptions",
    "6200": "Professional services: legal, accounting, consulting",
    "6300": "Travel and entertainment",
    "6400": "Facilities: rent, utilities, maintenance",
}

AUTO_POST_THRESHOLD = 0.85  # tune against a sample your controller has already coded


def code_invoice(invoice_text: str) -> dict:
    response = client.system_one(
        state=invoice_text,
        questions={
            "gl_account": Choice(
                instructions="Which general ledger account should this invoice be coded to",
                criteria=GL_ACCOUNTS,
            ),
        },
    )
    answer = response.answers["gl_account"]
    route = "auto_post" if answer.confidence >= AUTO_POST_THRESHOLD else "ap_review_queue"
    return {
        "gl_account": answer.choice,
        "confidence": answer.confidence,
        "route": route,
    }
from typesafe_sdk import Choice, TypeSafeClient

client = TypeSafeClient()

GL_ACCOUNTS = {
    "5100": "Cost of goods: raw materials and inventory",
    "6100": "Software and SaaS subscriptions",
    "6200": "Professional services: legal, accounting, consulting",
    "6300": "Travel and entertainment",
    "6400": "Facilities: rent, utilities, maintenance",
}

AUTO_POST_THRESHOLD = 0.85  # tune against a sample your controller has already coded


def code_invoice(invoice_text: str) -> dict:
    response = client.system_one(
        state=invoice_text,
        questions={
            "gl_account": Choice(
                instructions="Which general ledger account should this invoice be coded to",
                criteria=GL_ACCOUNTS,
            ),
        },
    )
    answer = response.answers["gl_account"]
    route = "auto_post" if answer.confidence >= AUTO_POST_THRESHOLD else "ap_review_queue"
    return {
        "gl_account": answer.choice,
        "confidence": answer.confidence,
        "route": route,
    }

Example 2: First-pass deal screening with weights you control

We recommend breaking complex judgments into small, atomic questions and combining them in code. For a deal team, that means the investment thesis lives in a few lines of Python, not buried in a prompt. When priorities change, you change a weight.

from typesafe_sdk import Noul, Score, TypeSafeClient

client = TypeSafeClient()

THREE_LEVELS = 2  # Score returns 0, 1 or 2 for a three-level rubric

QUESTIONS = {
    "recurring_revenue": Score(
        instructions="How much of the company's revenue is recurring or contracted",
        criteria=["Mostly project-based or transactional", "Mixed", "Mostly recurring or contracted"],
    ),
    "customer_diversification": Score(
        instructions="How diversified the customer base is",
        criteria=["A few customers dominate revenue", "Moderately concentrated", "Well diversified"],
    ),
    "margin_profile": Score(
        instructions="How the company's margins are trending",
        criteria=["Thin or declining", "Stable and in line with peers", "Expanding or above peers"],
    ),
    "founder_owned": Noul(
        instructions="The company is founder- or family-owned with no prior institutional investors",
    ),
}

WEIGHTS = {"recurring_revenue": 0.40, "customer_diversification": 0.35, "margin_profile": 0.25}


def screen_teaser(teaser_text: str) -> dict:
    answers = client.system_one(state=teaser_text, questions=QUESTIONS).answers
    fit_score = sum(w * answers[k].score / THREE_LEVELS for k, w in WEIGHTS.items())
    return {
        "fit_score": round(fit_score, 2),  # 0.0 to 1.0
        "founder_owned": answers["founder_owned"].noul > 0.5,
        "take_first_call": fit_score >= 0.6,
    }
from typesafe_sdk import Noul, Score, TypeSafeClient

client = TypeSafeClient()

THREE_LEVELS = 2  # Score returns 0, 1 or 2 for a three-level rubric

QUESTIONS = {
    "recurring_revenue": Score(
        instructions="How much of the company's revenue is recurring or contracted",
        criteria=["Mostly project-based or transactional", "Mixed", "Mostly recurring or contracted"],
    ),
    "customer_diversification": Score(
        instructions="How diversified the customer base is",
        criteria=["A few customers dominate revenue", "Moderately concentrated", "Well diversified"],
    ),
    "margin_profile": Score(
        instructions="How the company's margins are trending",
        criteria=["Thin or declining", "Stable and in line with peers", "Expanding or above peers"],
    ),
    "founder_owned": Noul(
        instructions="The company is founder- or family-owned with no prior institutional investors",
    ),
}

WEIGHTS = {"recurring_revenue": 0.40, "customer_diversification": 0.35, "margin_profile": 0.25}


def screen_teaser(teaser_text: str) -> dict:
    answers = client.system_one(state=teaser_text, questions=QUESTIONS).answers
    fit_score = sum(w * answers[k].score / THREE_LEVELS for k, w in WEIGHTS.items())
    return {
        "fit_score": round(fit_score, 2),  # 0.0 to 1.0
        "founder_owned": answers["founder_owned"].noul > 0.5,
        "take_first_call": fit_score >= 0.6,
    }
from typesafe_sdk import Noul, Score, TypeSafeClient

client = TypeSafeClient()

THREE_LEVELS = 2  # Score returns 0, 1 or 2 for a three-level rubric

QUESTIONS = {
    "recurring_revenue": Score(
        instructions="How much of the company's revenue is recurring or contracted",
        criteria=["Mostly project-based or transactional", "Mixed", "Mostly recurring or contracted"],
    ),
    "customer_diversification": Score(
        instructions="How diversified the customer base is",
        criteria=["A few customers dominate revenue", "Moderately concentrated", "Well diversified"],
    ),
    "margin_profile": Score(
        instructions="How the company's margins are trending",
        criteria=["Thin or declining", "Stable and in line with peers", "Expanding or above peers"],
    ),
    "founder_owned": Noul(
        instructions="The company is founder- or family-owned with no prior institutional investors",
    ),
}

WEIGHTS = {"recurring_revenue": 0.40, "customer_diversification": 0.35, "margin_profile": 0.25}


def screen_teaser(teaser_text: str) -> dict:
    answers = client.system_one(state=teaser_text, questions=QUESTIONS).answers
    fit_score = sum(w * answers[k].score / THREE_LEVELS for k, w in WEIGHTS.items())
    return {
        "fit_score": round(fit_score, 2),  # 0.0 to 1.0
        "founder_owned": answers["founder_owned"].noul > 0.5,
        "take_first_call": fit_score >= 0.6,
    }

What to watch before you commit

Jev is two weeks old and still in early access. It's promising, but treat it like any new vendor in a portfolio company stack.

  • Performance claims are vendor-reported. Published speed multiples range from 40x to 200x depending on the source and task. TypeSafe itself notes that its internal workflow results likely sit at the high end of what customers will see.

  • Limited transparency. TypeSafe has not published Jev's architecture, weights or a technical paper.

  • Known rough edges. TypeSafe documents the current version's weak spots publicly. Read that page before scoping a pilot.

  • Question design matters. Jev performs best on narrow, gut-check questions. Anything requiring extended reasoning should be broken into smaller questions or left to an LLM.

  • Data governance. Jev runs as a hosted API. Confirm data handling terms before sending customer, financial or employee data from a portfolio company.

The practical path: pick one high-volume workflow, benchmark Jev against your current approach on a labeled sample, and set confidence thresholds before anything goes live.

Put Jev to work in your portfolio

BaseForge Advisors helps middle-market private equity firms and their portfolio companies turn AI from pilots into measurable EBITDA impact. If you're weighing where a System One model fits in your stack, or want to pilot Jev on a live workflow, reach out to our team.