What Are System 1 Decision Models: Ending Generative Token Waste

September 22, 2026

What Are System 1 Decision Models: Ending Generative Token Waste

A conceptual and mathematical deep dive into prefill-only decision models. Why forcing generative LLMs into JSON structured outputs is an architectural flaw, and how Jev, Laya, and Kev establish a new standard.

·8 min read·Parsec AI Labs
← back to the blog

For the past four years, the artificial intelligence industry has normalized an engineering anti-pattern: deploying monolithic autoregressive transformers trained on next-token prediction to solve strictly taxonomic, deterministic problems.

When a software architecture needs to determine whether a customer support ticket indicates billing fraud, decide which engineering rotation should receive an alert, or verify whether an autonomous agent’s proposed bash command violates safety guardrails, invoking GPT-4o, Claude 3.5 Sonnet, or Llama 3 with Structured Outputs (forced JSON schemas) is the computational equivalent of firing up a jet engine to warm a cup of coffee.

In mid-September 2026, a structural conceptual shift crystallized with the arrival of System 1 decision models (prefill-only non-generative architectures), pioneered by projects including Jev (TypeSafe AI), Laya (Convai Innovations), and Kev (Jared Palmer). To understand why this represents a pivotal turning point, it is necessary to examine both the physical bottlenecks of generative decoding and the mathematics underlying this new paradigm.


The Cognitive Duality: System 1 vs. System 2

In his seminal work Thinking, Fast and Slow (2011), psychologist and Nobel laureate Daniel Kahneman formalized human cognition into two distinct modes of information processing:

  1. System 1: Fast, automatic, parallel, subconscious, and heuristic. It operates effortlessly without deliberate reflection. It is the cognitive mechanism that recognizes a familiar face across a crowded room, ducks an incoming projectile, or gauges vocal hostility in under 100 milliseconds.
  2. System 2: Slow, sequential, deliberative, conscious, and computationally expensive. It requires step-by-step attention. It is the mechanism deployed when multiplying three-digit numbers, mentally verifying a code snippet, or calculating chess tactics.

Between 2022 and 2025, frontier AI development produced a monumental System 2: dense reasoning models (such as OpenAI o1 or DeepSeek-R1) that deliberate sequentially via extended Chain-of-Thought reasoning.

However, software engineers made the critical mistake of forcing that System 2 to masquerade as System 1. Lacking an agile, non-generative counterpart, we embedded multi-billion-parameter text generators into every logical branch of our production applications.


Why Generative LLMs Break on Classification Tasks

Forcing an autoregressive language model to yield typed decisions via text generation introduces three severe physical bottlenecks:

1. The Memory Bandwidth Bottleneck of Autoregressive Decoding

The decoding phase of generative LLMs is fundamentally bound by memory bandwidth (memory bandwidth bound), not raw compute capacity (FLOPs). Generating each individual token requires streaming the full weight tensors of the model from GPU VRAM to the compute cores while incrementally updating a growing KV Cache:

Time per token ≈ (2 × Parameters) / (Memory Bandwidth of VRAM)

When a 70-billion-parameter model generates 40 tokens of JSON syntax ({"classification": "urgent", "score": 0.95}), the GPU hardware must perform 40 sequential passes across memory. This establishes a physical latency floor of 400ms to 2,500ms, which cannot be eliminated across standard client-server APIs.

2. State Bloat and VRAM Thrashing

In modern distributed serving systems (such as vLLM or TGI), holding intermediate KV cache allocations open while the model decodes purely syntactic formatting characters consumes memory headroom that could otherwise serve concurrent requests.

3. Non-Deterministic Failure Modes

Even when using constrained grammars (GBNF) or token-masking heuristics to enforce JSON schema compliance, autoregressive models remain vulnerable to subtle semantic drift, variable response times, and unpredictable latencies dependent on token length.


The Anatomy of a Decision Model: The Prefill-Only Approach

A decision model completely excises the autoregressive decoding loop. There is no previous token, no next token, and no generation buffer.

Instead, the system consumes a typed pair: a State (S) and a collection of Typed Questions (Q_1, Q_2, ... Q_m). The entire operation executes in a single forward pass (prefill phase):

Input [State S + Typed Question Q]
              │
              ▼
   ┌──────────────────────┐
   │ Transformer Backbone │ ──► O(1) Parallel Prefill Phase
   │  (Encoder or Frozen) │     (Dense Attention Tensor)
   └──────────────────────┘
              │
              ▼ Latent Vector h_L
   ┌──────────────────────┐
   │ Decision Readout     │
   │ Projection Heads     │
   └──────────────────────┘
      │           │          │
      ▼           ▼          ▼
    Noul       Choice      Score
   p ∈ [0,1]  Softmax(k)  Ordinal Scale
  1. Dense Parallel Processing: The tokens comprising both the state payload and the query are ingested simultaneously through attention blocks via optimized kernels (FlashAttention / Tensor Cores), completing context ingestion in a single O(1) step.
  2. Latent Representation Extraction (h_L): The model extracts the hidden representation from the final sequence position or through learned attention pooling.
  3. Linear Projection Heads: Instead of computing logits over a 128,000-token vocabulary, h_L is projected directly through compact, task-specific linear matrices:
  • Noul (Boolean Probability): Scalar projection through a sigmoid activation:
p = sigmoid(W_noul * h_L + b) = 1 / (1 + exp(-(W_noul * h_L + b)))
  • Choice (Categorical Classification): k-dimensional projection normalized by softmax:
P(y = i) = exp(W_i * h_L) / sum_j(exp(W_j * h_L))
  • Score (Bounded Ordinal): Projection into a calibrated numeric interval or ordinal logistic regression.

The output is delivered as an immediate mathematical probability structure in 25 to 60 milliseconds.


The Mathematics of Calibration: Beyond Naive Classifiers

In production software pipelines, a probability emitted by an AI model is useless if it is not statistically calibrated.

Generative LLMs suffer from chronic overconfidence: when an LLM claims 99% certainty, empirical tests show it is often accurate in fewer than 80% of edge cases. Standard cross-entropy objective functions and human-preference RLHF tune models to sound authoritative rather than mathematically precise.

A true decision model is trained and evaluated using Strictly Proper Scoring Rules, such as the Brier Score:

Brier Score (BS) = (1 / N) * sum_{t=1..N} (f_t - o_t)^2

Where f_t represents the model’s reported probability forecast and o_t represents the ground-truth binary outcome (0 or 1). A strictly proper scoring rule ensures that the expected reward is maximized if and only if the model reports its exact Bayesian belief.

Consequently, when a calibrated decision model reports a confidence probability of 0.85, exactly 85% of historical samples within that bin prove true in practice.


The Pioneer Ecosystem: Jev, Laya, and Kev

This conceptual architecture has been realized through three complementary implementations:

Project Format Backbone Architecture Primary Design Philosophy
Jev Proprietary (Cloud API) Native non-autoregressive network Developed by TypeSafe AI ($40M seed led by DCVC). Engineered for ultra-high-throughput enterprise systems with simultaneous multi-question parallel evaluation.
Laya Open Source (Apache 2.0) ModernBERT-large (421M) / mmBERT (322M) Developed by Convai Innovations. Pure encoder-only architecture trained with RLCD for mathematical calibration and on-premise data sovereignty.
Kev Open Source (Apache 2.0) Qwen 3.5 (0.8B, 4B, 9B) + LoRA Engineered by Jared Palmer (Cognition / Devin). Proves that open decoder LLMs can be frozen into prefill-only local judgment engines compatible with TypeSafe’s API.

Design Pattern: Dual-Tier Agent Architectures

The architectural role of System 1 decision models is not to displace deep reasoning LLMs, but to serve as an indispensable front-line gatekeeper.

In production agentic architectures, 85% to 95% of incoming user intents or events are standard operations (factual queries, standard routing, safe command invocations). Deploying a reasoning model to evaluate these tasks represents significant infrastructural waste.

Incoming User Request / Event
              │
              ▼
┌──────────────────────────────────────┐
│  System 1 Gatekeeper (Jev/Laya/Kev)  │  Latency: ~30ms | Cost: ~$0.0001
│  - Is this a complex inquiry?        │
│  - Does this violate policy rules?   │
│  - Does this require tool execution? │
└──────────────────────────────────────┘
      │                           │
  [p > 0.95: High Confidence] [p < 0.95: Ambiguous / Complex]
      │                           │
      ▼                           ▼
┌───────────────────────┐   ┌────────────────────────────────┐
│  Fast Path Execution  │   │  System 2 Deliberative Core    │
│  (Docs / Cache / SQL) │   │  (DeepSeek-R1 / Claude 3.5 / o1│
│  Total Latency <60ms  │   │  Sequential reasoning chains   │
└───────────────────────┘   └────────────────────────────────┘

By interposing a System 1 gatekeeper, engineering teams eliminate up to 90% of their inferencing bills, neutralize prompt injection risks in milliseconds, and deliver instant sub-100ms user experiences.

System 1 decision models are not a superficial variation of conversational LLMs. They are the missing foundational primitive required to build predictable, low-latency, and cost-effective AI software systems.


References

Related notes