For the past four years, the artificial intelligence industry has normalized an engineering anti-pattern: deploying monolithic autoregressive transformers trained on next-token prediction to solve strictly taxonomic, deterministic problems.
When a software architecture needs to determine whether a customer support ticket indicates billing fraud, decide which engineering rotation should receive an alert, or verify whether an autonomous agent’s proposed bash command violates safety guardrails, invoking GPT-4o, Claude 3.5 Sonnet, or Llama 3 with Structured Outputs (forced JSON schemas) is the computational equivalent of firing up a jet engine to warm a cup of coffee.
In mid-September 2026, a structural conceptual shift crystallized with the arrival of System 1 decision models (prefill-only non-generative architectures), pioneered by projects including Jev (TypeSafe AI), Laya (Convai Innovations), and Kev (Jared Palmer). To understand why this represents a pivotal turning point, it is necessary to examine both the physical bottlenecks of generative decoding and the mathematics underlying this new paradigm.
The Cognitive Duality: System 1 vs. System 2
In his seminal work Thinking, Fast and Slow (2011), psychologist and Nobel laureate Daniel Kahneman formalized human cognition into two distinct modes of information processing:
- System 1: Fast, automatic, parallel, subconscious, and heuristic. It operates effortlessly without deliberate reflection. It is the cognitive mechanism that recognizes a familiar face across a crowded room, ducks an incoming projectile, or gauges vocal hostility in under 100 milliseconds.
- System 2: Slow, sequential, deliberative, conscious, and computationally expensive. It requires step-by-step attention. It is the mechanism deployed when multiplying three-digit numbers, mentally verifying a code snippet, or calculating chess tactics.
Between 2022 and 2025, frontier AI development produced a monumental System 2: dense reasoning models (such as OpenAI o1 or DeepSeek-R1) that deliberate sequentially via extended Chain-of-Thought reasoning.
However, software engineers made the critical mistake of forcing that System 2 to masquerade as System 1. Lacking an agile, non-generative counterpart, we embedded multi-billion-parameter text generators into every logical branch of our production applications.
Why Generative LLMs Break on Classification Tasks
Forcing an autoregressive language model to yield typed decisions via text generation introduces three severe physical bottlenecks:
1. The Memory Bandwidth Bottleneck of Autoregressive Decoding
The decoding phase of generative LLMs is fundamentally bound by memory bandwidth (memory bandwidth bound), not raw compute capacity (FLOPs). Generating each individual token requires streaming the full weight tensors of the model from GPU VRAM to the compute cores while incrementally updating a growing KV Cache:
Time per token ≈ (2 × Parameters) / (Memory Bandwidth of VRAM)
When a 70-billion-parameter model generates 40 tokens of JSON syntax ({"classification": "urgent", "score": 0.95}), the GPU hardware must perform 40 sequential passes across memory. This establishes a physical latency floor of 400ms to 2,500ms, which cannot be eliminated across standard client-server APIs.
2. State Bloat and VRAM Thrashing
In modern distributed serving systems (such as vLLM or TGI), holding intermediate KV cache allocations open while the model decodes purely syntactic formatting characters consumes memory headroom that could otherwise serve concurrent requests.
3. Non-Deterministic Failure Modes
Even when using constrained grammars (GBNF) or token-masking heuristics to enforce JSON schema compliance, autoregressive models remain vulnerable to subtle semantic drift, variable response times, and unpredictable latencies dependent on token length.
The Anatomy of a Decision Model: The Prefill-Only Approach
A decision model completely excises the autoregressive decoding loop. There is no previous token, no next token, and no generation buffer.
Instead, the system consumes a typed pair: a State (S) and a collection of Typed Questions (Q_1, Q_2, ... Q_m). The entire operation executes in a single forward pass (prefill phase):
Input [State S + Typed Question Q]
│
▼
┌──────────────────────┐
│ Transformer Backbone │ ──► O(1) Parallel Prefill Phase
│ (Encoder or Frozen) │ (Dense Attention Tensor)
└──────────────────────┘
│
▼ Latent Vector h_L
┌──────────────────────┐
│ Decision Readout │
│ Projection Heads │
└──────────────────────┘
│ │ │
▼ ▼ ▼
Noul Choice Score
p ∈ [0,1] Softmax(k) Ordinal Scale
- Dense Parallel Processing: The tokens comprising both the state payload and the query are ingested simultaneously through attention blocks via optimized kernels (FlashAttention / Tensor Cores), completing context ingestion in a single O(1) step.
- Latent Representation Extraction (
h_L): The model extracts the hidden representation from the final sequence position or through learned attention pooling. - Linear Projection Heads: Instead of computing logits over a 128,000-token vocabulary,
h_Lis projected directly through compact, task-specific linear matrices:
Noul(Boolean Probability): Scalar projection through a sigmoid activation:
p = sigmoid(W_noul * h_L + b) = 1 / (1 + exp(-(W_noul * h_L + b)))
Choice(Categorical Classification):k-dimensional projection normalized by softmax:
P(y = i) = exp(W_i * h_L) / sum_j(exp(W_j * h_L))
Score(Bounded Ordinal): Projection into a calibrated numeric interval or ordinal logistic regression.
The output is delivered as an immediate mathematical probability structure in 25 to 60 milliseconds.
The Mathematics of Calibration: Beyond Naive Classifiers
In production software pipelines, a probability emitted by an AI model is useless if it is not statistically calibrated.
Generative LLMs suffer from chronic overconfidence: when an LLM claims 99% certainty, empirical tests show it is often accurate in fewer than 80% of edge cases. Standard cross-entropy objective functions and human-preference RLHF tune models to sound authoritative rather than mathematically precise.
A true decision model is trained and evaluated using Strictly Proper Scoring Rules, such as the Brier Score:
Brier Score (BS) = (1 / N) * sum_{t=1..N} (f_t - o_t)^2
Where f_t represents the model’s reported probability forecast and o_t represents the ground-truth binary outcome (0 or 1). A strictly proper scoring rule ensures that the expected reward is maximized if and only if the model reports its exact Bayesian belief.
Consequently, when a calibrated decision model reports a confidence probability of 0.85, exactly 85% of historical samples within that bin prove true in practice.
The Pioneer Ecosystem: Jev, Laya, and Kev
This conceptual architecture has been realized through three complementary implementations:
| Project | Format | Backbone Architecture | Primary Design Philosophy |
|---|---|---|---|
| Jev | Proprietary (Cloud API) | Native non-autoregressive network | Developed by TypeSafe AI ($40M seed led by DCVC). Engineered for ultra-high-throughput enterprise systems with simultaneous multi-question parallel evaluation. |
| Laya | Open Source (Apache 2.0) | ModernBERT-large (421M) / mmBERT (322M) | Developed by Convai Innovations. Pure encoder-only architecture trained with RLCD for mathematical calibration and on-premise data sovereignty. |
| Kev | Open Source (Apache 2.0) | Qwen 3.5 (0.8B, 4B, 9B) + LoRA | Engineered by Jared Palmer (Cognition / Devin). Proves that open decoder LLMs can be frozen into prefill-only local judgment engines compatible with TypeSafe’s API. |
Design Pattern: Dual-Tier Agent Architectures
The architectural role of System 1 decision models is not to displace deep reasoning LLMs, but to serve as an indispensable front-line gatekeeper.
In production agentic architectures, 85% to 95% of incoming user intents or events are standard operations (factual queries, standard routing, safe command invocations). Deploying a reasoning model to evaluate these tasks represents significant infrastructural waste.
Incoming User Request / Event
│
▼
┌──────────────────────────────────────┐
│ System 1 Gatekeeper (Jev/Laya/Kev) │ Latency: ~30ms | Cost: ~$0.0001
│ - Is this a complex inquiry? │
│ - Does this violate policy rules? │
│ - Does this require tool execution? │
└──────────────────────────────────────┘
│ │
[p > 0.95: High Confidence] [p < 0.95: Ambiguous / Complex]
│ │
▼ ▼
┌───────────────────────┐ ┌────────────────────────────────┐
│ Fast Path Execution │ │ System 2 Deliberative Core │
│ (Docs / Cache / SQL) │ │ (DeepSeek-R1 / Claude 3.5 / o1│
│ Total Latency <60ms │ │ Sequential reasoning chains │
└───────────────────────┘ └────────────────────────────────┘
By interposing a System 1 gatekeeper, engineering teams eliminate up to 90% of their inferencing bills, neutralize prompt injection risks in milliseconds, and deliver instant sub-100ms user experiences.
System 1 decision models are not a superficial variation of conversational LLMs. They are the missing foundational primitive required to build predictable, low-latency, and cost-effective AI software systems.
References
- Thinking, Fast and Slow (Daniel Kahneman) — Foundational cognitive psychology theory on dual-process architectures. Date: 2011-10-25. Type: Reference book.
- Strictly Proper Scoring Rules, Prediction, and Estimation — Gneiting & Raftery. Mathematical foundations of probability calibration and Brier score minimization. Date: 2007-03-01. Type: Research paper.
- TypeSafe AI: Introducing Jev — Technical release announcement regarding Machine-Native Intelligence and System One decision engines. Date: 2026-09-15. Type: Official documentation.
- NandhaKishorM/laya: Open System 1 Decision Model — Codebase and architecture specifications for Laya across ModernBERT and mmBERT backbones. Date: 2026-09-18. Type: Open source repository.
- jaredpalmer/kev: Drop-in Local Decision Models — Prefill-only implementation built atop Qwen 3.5 using pointer readout heads and LoRA adapters. Date: 2026-09-19. Type: Open source repository.
- ModernBERT: Modernizing BERT for Fast and Long-Context Inference — Official paper detailing the ModernBERT encoder architecture. Date: 2024-12-18. Type: Research paper.


