Background
The evolution of artificial intelligence has been heavily characterized by the tension between unbounded generative capabilities and the strict requirements of programmatic execution. Autoregressive language models excel at free-text generation, sequentially predicting tokens based on immense contextual histories. However, when deployed as decision engines or control-flow mechanisms, this sequential architecture introduces massive overhead, unpredictable latency spikes, and structural fragility. The fundamental mismatch lies in using an architecture designed for open-ended semantic formulation to perform highly constrained, categorical decision-making. Software engineering has required structured, typed decisions for decades, and the machine learning field has continuously attempted to bridge this gap through various architectures and paradigms well before the advent of frontier language models.
The landscape before Jev
The necessity for programmatic decision-making in machine learning is not novel. The landscape of structured evaluation is vast and can be categorically divided into distinct technique families, each addressing the core problems of latency, cost, calibration, and output structure. These families emerged chronologically as the field evolved, but they are best understood through their architectural intent. For decades, engineers have needed to classify data, route requests based on content, and extract specific entities from unstructured text. The journey from highly rigid rule-based systems to the flexible but unpredictable frontier language models of today highlights a continuous trade-off between the rigidity of typed execution and the intelligence required to parse ambiguous human language.
Classical decision systems Classical decision systems encompass rule engines, logistic regression models, and gradient-boosted classifiers such as XGBoost or LightGBM. Emerging in the late 1990s and solidifying through the 2010s, these systems turn input into a decision by mapping fixed-dimensional feature vectors to hyperplanes or decision trees, outputting strict categorical labels [22]. Crucially, they excel at returning calibrated probabilities, where the output score genuinely reflects the empirical likelihood of correctness. Methods like Platt scaling (fitting a logistic regression model to the classifier’s scores) and isotonic regression (a non-parametric approach fitting a non-decreasing step function) have long been used to map raw margins to true empirical probabilities [1, 22]. These models are characterized by sub-millisecond latency and negligible inference costs, making them the default for high-throughput routing. However, their primary failure mode is their extreme brittleness to unstructured data; they require extensive feature engineering and massive amounts of labeled training data to generalize across varying input distributions. They give typed outputs natively and multiple decisions per pass if configured as multi-target regressors. Because their fundamental operations involve simple matrix multiplications and thresholding rather than deep sequential attention mechanisms, they can easily process millions of requests per second on standard hardware, making them the gold standard for latency-critical applications where the input domain is strictly bounded.
Fine-tuned encoder classifiers
As deep learning matured, fine-tuned encoder classifiers became the standard for unstructured text evaluation. Emerging prominently with BERT in 2018, these architectures utilize bidirectional self-attention to process entire input contexts simultaneously. A classification head (typically a linear layer followed by a softmax activation) is placed over the [CLS] token representation, mapping the dense semantic vector directly to a typed categorical output. These models execute a single forward pass, resulting in latencies ranging from 10ms to 50ms depending on sequence length and hardware. They can easily produce multiple decisions per pass by incorporating multi-task or multi-head architectures over the same base encoder. While highly efficient, standard encoder classifiers are notorious for being poorly calibrated; they often yield overconfident probabilities even when incorrect [1]. Modern iterations, such as SetFit, apply contrastive learning and distillation techniques to achieve high accuracy with minimal training examples, yet the fundamental requirement for task-specific fine-tuning remains [12]. The contrastive learning phase helps the model rapidly differentiate between classes even with only a handful of examples per class, drastically reducing the data acquisition burden that plagued classical systems. Nevertheless, deploying these models still necessitates managing discrete weights and endpoints for every unique classification task an organization wishes to perform, creating significant MLOps overhead as the number of required decisions scales.
Zero-shot classification To eliminate the dependency on task-specific training data, researchers developed zero-shot classification methodologies. One highly successful approach frames classification as a Natural Language Inference (NLI) or entailment problem [2]. By constructing a premise (the unstructured state) and a hypothesis (the proposed label), the model calculates the probability of entailment. This emerged around 2019 and allowed for typed outputs without bespoke training. More recently, architectures like GLiNER and GLiClass have introduced single-pass multi-label zero-shot models that utilize bidirectional transformers to directly map text spans to an arbitrary set of provided labels [13]. These models offer pseudo-calibrated probabilities based on entailment scores and can evaluate multiple questions simultaneously. Their latency profile is slightly higher than fine-tuned encoders due to the complexity of pairwise entailment computations, but they eliminate the data acquisition bottleneck. Their known failure mode involves struggling with highly nuanced or domain-specific logic that falls outside their pre-training distribution. Because they infer labels by determining if a text entails a specific phrase, they often fail when the difference between two categories hinges on extremely subtle contextual clues or domain-specific jargon that was not adequately represented in the broad entailment datasets they were trained upon.
The System 1 / System 2 framing in AI The conceptual foundation for differentiating rapid, intuitive decision-making from deliberate, sequential reasoning was formalized by Kahneman and Tversky and subsequently adapted into AI research [8]. “Thinking Fast and Slow in AI” maps these cognitive systems to neural architectures: System 1 represents fast, parallel, pattern-matching execution (such as a forward pass in a classifier), while System 2 represents slower, sequential, algorithmic formulation (such as chain-of-thought autoregressive generation). This framing is critical for understanding the landscape. For years, the AI industry conflated these systems, utilizing System 2 architectures (autoregressive LLMs) to perform System 1 tasks (classification, routing, and state evaluation). The inefficiencies of this conflation are exactly what the subsequent generations of AI tooling and models have attempted to rectify. TypeSafe AI utilizes this exact System 1 framing to differentiate its specialized architecture from generalized LLMs, though the underlying distinction between classification and generation is foundational.
Current State
The current state of structured evaluation is defined by the tension between adapting existing autoregressive models via heavy programmatic guardrails and developing entirely new architectures optimized specifically for the task. TypeSafe AI has introduced Jev, which abandons the adaptation approach entirely, positioning itself explicitly as a “System One” model designed for frontier-intelligence function calling.
What Jev does
Jev is an AI model designed to evaluate unstructured state against a set of predeclared, typed questions, returning strict probabilistic decisions without engaging in free-text string generation. According to TypeSafe AI’s release documentation, the Jev interface operates on a strictly constrained input-output paradigm. It takes an unstructured state context and a predefined schema of possible outputs [23]. Because Jev calculates its outputs in parallel rather than sequentially generating tokens, it mathematically eliminates type errors and hallucinations; the model can only return values that correspond to the predeclared structural keys provided in its query [23]. This is a massive departure from standard LLMs that generate strings that happen to look like JSON.
The model’s internal architecture relies entirely on parallel sampling, computing all requested probabilities in a single query. TypeSafe AI claims this architectural shift results in severe reductions in latency, with end-to-end response times ranging from 70ms to 500ms, which they benchmark as 40x to 200x faster than traditional LLMs executing comparable structured tasks [23]. The economic model shifts correspondingly. While autoregressive models charge heavily for output generation (where the cost is dominated by sequential memory bandwidth limits and autoregressive decoding overhead), Jev prices input tokens at $0.042 per million and makes output tokens completely free [23].
To ensure the probabilities returned are epistemically honest, TypeSafe trained Jev using a novel methodology termed Reinforcement Learning for Calibrated Decisions (RLCD). Instead of optimizing for human preference (RLHF) or verifiable programmatic rewards (RLVR), RLCD heavily penalizes overconfidence on incorrect answers and underconfidence on correct ones, attempting to align the model’s output confidence directly with its empirical accuracy. Every answer emitted by Jev includes these calibrated confidence scores [23]. TypeSafe’s benchmark setup evaluated Jev across four complex, real-world workflows involving high-cardinality routing and discrete branching logic. They used the average predictions of frontier models (GPT-6 Astra and Fable 5.1) as the reference probabilities to test against. Jev owned the Pareto frontier in these evaluations, demonstrating comparable intelligence to GPT-5.6 Terra on System 1 shaped queries while maintaining its massive speed and cost advantages [23].
How Jev differs
Jev’s approach combines specific elements from prior architectural families while introducing novel synthesis and packaging. The table below compares Jev against the fundamental technique families evaluated in the background section.
| Technique Family | Native Typed Outputs | Multi-Question per Call | Calibrated Output | Needs Fine-Tuning Data? | Latency Class |
|---|---|---|---|---|---|
| Classical Decision Systems | Yes | Yes (Multi-target) | Yes (Platt/Isotonic) | Yes (Extensive) | <1ms |
| Fine-Tuned Encoders | Yes | Yes (Multi-head) | No (Often overconfident) | Yes (Moderate) | 10ms - 50ms |
| Zero-Shot (NLI/GLiClass) | Yes | Yes | Pseudo-calibrated | No | 50ms - 100ms |
| TypeSafe Jev | Yes | Yes | Yes (via RLCD) | No | 70ms - 500ms |
Jev essentially combines the zero-shot intelligence and lack of training requirements found in modern foundation models with the parallel execution speed of classical classifiers, while newly claiming RLCD-driven intrinsic calibration at the frontier intelligence level. Its architectural novelty lies in training a massive parameter model exclusively via RLCD for parallel discrete classification rather than next-token prediction. Its packaging novelty lies in the interface, which abstracts away the complexities of logit extraction and presents a clean, type-safe API where output tokens are economically free [23]. By shifting the computational burden entirely to the parallel evaluation of the state against the predeclared schema, Jev fundamentally alters the unit economics of using frontier intelligence for structured tasks. The user pays only to contextualize the model with the unstructured state, and the subsequent extraction of probabilities for the requested types incurs almost zero marginal cost due to the lack of sequential decoding steps.
Alternatives available today
While Jev introduces a specialized architecture, the problem of structured evaluation can be solved today using a variety of existing methodologies. The landscape of available alternatives is broad, ranging from API-level features on frontier models to sophisticated open-source orchestration frameworks.
| Alternative | Needs Training Data? | Multi-Question per Call | Calibrated Output | Latency Class | Cost Class | Self-Hostable? | Maturity |
|---|---|---|---|---|---|---|---|
| Frontier LLM Structured Outputs (OpenAI) | No | Yes (via JSON schema) | No (Hidden logits) | >1000ms | High | No | High |
| Typed Frameworks (DSPy, Instructor, BAML) | No (Prompt engineered) | Yes (via Pydantic) | No | >1000ms | High/Medium | Yes (with open weights) | High |
| Guard/Judge Models (Llama Guard) | No (Pre-trained) | No (Single category) | Pseudo-calibrated | 100ms - 500ms | Low | Yes | Medium |
| Zero-Shot Models (GLiClass, NLI) | No | Yes | Pseudo-calibrated | 50ms - 100ms | Low | Yes | High |
| Fine-Tuned Encoders (SetFit) | Yes (Few-shot) | Yes | No | 10ms - 50ms | Very Low | Yes | High |
| LLM Routers (RouteLLM) | Yes (Preference data) | No (Single route) | Yes (Thresholded) | 10ms - 50ms | Low | Yes | Medium |
Frontier LLM Structured Outputs: Providers like OpenAI offer native Structured Outputs and JSON mode. This utilizes constrained decoding at the API level to guarantee schema adherence [30]. It is the better choice than Jev when the task inherently requires deep semantic formulation or chain-of-thought reasoning before arriving at a structured decision, as Jev’s parallel nature prevents sequential reasoning traces.
Typed Frameworks: Orchestration layers like Instructor [25], DSPy [24], Outlines [26], and BAML [27] wrap standard autoregressive models. Instructor patches the OpenAI API to return Pydantic models via function calling [31]. DSPy compiles declarative signatures into optimized prompts [16]. Outlines utilizes finite state machines to enforce regex and grammar constraints during decoding [26]. BAML implements a specialized parser to extract structured types from messy LLM outputs [27]. These are the better choice than Jev when deploying to existing, self-hosted open-weight models where vendor lock-in to an API like TypeSafe is unacceptable.
Guard and Judge Models: Single-output transformer judges like Llama Guard [10] and ShieldGemma [11] are fine-tuned to return specific safety categories. They are the better choice than Jev for highly specific content moderation tasks where the underlying model has already been extensively aligned to human safety preferences and a specialized safety taxonomy.
Zero-Shot Models and Fine-Tuned Encoders: When latency requirements are strictly under 50ms and the semantic complexity of the task is narrow, fine-tuning a small BERT-style model via SetFit [12] or utilizing an NLI zero-shot classifier [2] remains the superior choice due to their negligible deployment costs and deterministic execution.
LLM Routers: Frameworks like RouteLLM utilize preference data to dynamically route queries between models of varying sizes based on difficulty [7]. They are the better choice than Jev when the objective is specifically to optimize cost across a cascade of different conversational LLMs rather than extracting state.
Future Outlook
The introduction of dedicated System One models like Jev validates a broader historical trend in machine learning research: the continuous attempt to constrain, calibrate, and accelerate large foundation models. Understanding Jev requires contextualizing it against the vast body of research that anticipated its core mechanisms.
Research with similar ideas
Scoring predefined answers with an LM instead of generating The idea of bypassing free-text generation by scoring predefined answers has deep roots in the literature. Cloze models and pattern-exploiting training (PET) formulate classification tasks as fill-in-the-blank problems, evaluating the exact probability of specific label tokens [3]. Instead of generating tokens, the system calculates the likelihood of predefined answers given the context. Research such as “Calibrate Before Use” demonstrated that language models exhibit severe contextual biases (e.g., favoring certain labels due to prompt structure) and proposed contextual calibration—an affine transformation of the output probabilities—to achieve robust few-shot performance [4]. By extracting exact log probabilities for restricted label tokens, researchers effectively turned autoregressive models into parallel evaluators for a single token, heavily anticipating Jev’s parallel schema evaluation.
Single-output transformer judges
The use of transformers as discrete evaluators rather than generators is widespread across the modern AI ecosystem. The LLM-as-a-judge paradigm leverages strong, highly capable foundation models to evaluate the unstructured text outputs of weaker models, essentially forcing the judge model to act as a discrete scoring function rather than a conversational agent [9]. Techniques like G-Eval utilize probability-weighted scoring, where instead of taking the single generated integer, the system calculates the expected value across the probabilities of all valid score tokens, creating a more continuous and calibrated signal [15]. Safety classifiers, such as Llama Guard and ShieldGemma, implement this by forcing the model to output specific typed categories (e.g., O1: Violence) based on predefined taxonomies [10, 11]. These approaches validate the utility of large parameter counts for singular, typed decisions.
Typed LLM interfaces and constrained decoding
To force autoregressive models to respect programmatic schemas, the community developed extensive tooling. Frameworks like Outlines and LMQL integrate directly with the decoding process, altering the logits at each step based on a finite state machine derived from a JSON schema or grammar [26]. While this guarantees type safety (a claim Jev mirrors), it incurs the heavy computational penalty of maintaining the autoregressive sequence generation [26]. DSPy abstracts the prompt engineering entirely, allowing developers to define declarative signatures (e.g., context, question -> answer) and automatically compiling them into optimal LLM calls [16, 24]. Microsoft’s TypeChat relies on prompting models with TypeScript interfaces and utilizing a rigorous validation loop to handle type errors [28, 29]. BoundaryML’s BAML parses highly unstructured string outputs into rigorous types, handling the reality that LLMs often wrap JSON in conversational markdown [27]. All these frameworks exist to solve the exact problem Jev solves natively.
LLM confidence and abstention TypeSafe’s RLCD training specifically optimizes for calibrated probabilities. This mirrors extensive research into whether language models “know what they know.” Studies have shown that while raw probabilities of label tokens can be well-calibrated, verbalized confidence (asking the model to generate its confidence score) is highly unreliable [5]. To address this, the field has explored conformal prediction for LLMs, a statistical framework that constructs prediction sets guaranteeing a specific coverage probability [18]. By generating a set of possible answers instead of a single overconfident one, models can effectively express uncertainty and abstain when necessary [18]. Jev’s core value proposition of epistemically honest probabilities builds directly on this lineage of calibration and uncertainty quantification research.
Cost and latency reduction for decisions The economic motivation behind Jev—offering massive cost reductions for structured evaluations—is preceded by cascade architectures. FrugalGPT demonstrated that by cascading queries from cheaper, smaller models to larger models only when the smaller model’s confidence falls below a learned threshold, users can achieve frontier performance at a fraction of the cost [6]. RouteLLM expanded on this by using preference data to train routers capable of directing traffic based on semantic complexity [7, 19]. These techniques rely on fast, calibrated decision-making to function; Jev effectively productizes the underlying requirement of these cascade systems into a single API.
Parallel generation Finally, while TypeSafe frames Jev around System One functionality, its underlying mechanics are deeply indebted to non-autoregressive (NAR) and parallel decoding research. Originally developed for neural machine translation to break the sequential bottleneck, NAR models attempt to predict all target tokens simultaneously [14]. Techniques like Mask-Predict utilize conditional masked language models to iteratively refine parallel predictions [20]. More recently, speculative decoding allows a smaller draft model to rapidly generate a sequence of tokens in parallel, which are then verified in a single forward pass by the larger target model [17, 21]. While these techniques attempt to retain the flexibility of unbounded string generation, Jev proves that by sacrificing that flexibility and restricting parallel execution exclusively to predefined schemas, models can achieve extreme latency reductions without sacrificing intelligence.
The future of automation likely involves a symbiotic relationship between these architectures. Systems like Jev will operate as the highly reliable, rapid control-flow mechanisms—the “smart if-statements” of the software stack—while traditional autoregressive models will be reserved strictly for tasks that inherently demand unbounded semantic formulation and sequential reasoning.
References
- Guo et al. “On Calibration of Modern Neural Networks.” arXiv:1706.04599 — https://arxiv.org/abs/1706.04599
- Yin et al. “Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach.” arXiv:1909.00161 — https://arxiv.org/abs/1909.00161
- Schick and Schütze. “Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference.” arXiv:2001.07676 — https://arxiv.org/abs/2001.07676
- Zhao et al. “Calibrate Before Use: Improving Few-Shot Performance of Language Models.” arXiv:2102.09690 — https://arxiv.org/abs/2102.09690
- Kadavath et al. “Language Models (Mostly) Know What They Know.” arXiv:2207.05221 — https://arxiv.org/abs/2207.05221
- Chen et al. “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.” arXiv:2305.05176 — https://arxiv.org/abs/2305.05176
- Hwang et al. “RouteLLM: Learning to Route LLMs with Preference Data.” arXiv:2406.18665 — https://arxiv.org/abs/2406.18665
- Booch et al. “Thinking Fast and Slow in AI.” arXiv:2010.06002 — https://arxiv.org/abs/2010.06002
- Zheng et al. “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.” arXiv:2306.05685 — https://arxiv.org/abs/2306.05685
- Inan et al. “Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.” arXiv:2312.06674 — https://arxiv.org/abs/2312.06674
- Team et al. “ShieldGemma: Generative AI Content Moderation Based on Gemma.” arXiv:2407.21772 — https://arxiv.org/abs/2407.21772
- Tunstall et al. “Efficient Few-Shot Learning Without Prompts.” arXiv:2209.11055 — https://arxiv.org/abs/2209.11055
- Zaratiana et al. “GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer.” arXiv:2311.08526 — https://arxiv.org/abs/2311.08526
- Gu et al. “Non-Autoregressive Neural Machine Translation.” arXiv:1711.02281 — https://arxiv.org/abs/1711.02281
- Liu et al. “G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.” arXiv:2303.16634 — https://arxiv.org/abs/2303.16634
- Khattab et al. “DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.” arXiv:2310.03714 — https://arxiv.org/abs/2310.03714
- Leviathan et al. “Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.” arXiv:2203.16487 — https://arxiv.org/abs/2203.16487
- Qu et al. “Domain-Shift-Aware Conformal Prediction for Large Language Models.” arXiv:2510.05566 — https://arxiv.org/abs/2510.05566
- Chen et al. “FrugalML: How to Use ML Prediction APIs More Accurately and Cheaply.” arXiv:2006.07512 — https://arxiv.org/abs/2006.07512
- Ghazvininejad et al. “Mask-Predict: Parallel Decoding of Conditional Masked Language Models.” arXiv:1904.09324 — https://arxiv.org/abs/1904.09324
- Miao et al. “Online Speculative Decoding.” arXiv:2310.07177 — https://arxiv.org/abs/2310.07177
- Gibbs et al. “Online Platt Scaling with Calibeating.” arXiv:2305.00070 — https://arxiv.org/abs/2305.00070
- TypeSafe AI. “Introducing System One Models & Jev.” 2026. https://typesafe.ai/blog/introducing-system-one-models-and-jev
- DSPy Contributors. “DSPy Repository.” 2024. https://github.com/stanfordnlp/dspy
- Instructor Contributors. “Instructor.” 2024. https://github.com/jxnl/instructor
- Outlines Contributors. “Outlines Repository.” 2024. https://github.com/outlines-dev/outlines
- BAML Contributors. “BoundaryML BAML.” 2024. https://github.com/boundaryml/baml
- Microsoft. “TypeChat Introduction.” 2024. https://microsoft.github.io/TypeChat/docs/introduction/
- Microsoft. “TypeChat FAQ.” 2024. https://microsoft.github.io/TypeChat/docs/faq/
- OpenAI. “Structured Outputs Guide.” 2024. https://platform.openai.com/docs/guides/structured-outputs
- Instructor Contributors. “Instructor Patching Concepts.” 2024. https://python.useinstructor.com/concepts/patching/