01 logo

A Random Number Generator Beat a Hyped AI Model at a Maze. Here’s Why AI Builders Should Pay Attention.

Jev ran 2,306 steps and got trapped in a corner. A Rust server cleared the maze in under 300ms. The lesson: System 1 models make fast decisions. They do not plan.

By JinPublished 9 days ago • 5 min read

A random-number server with Jev’s API format is running.

Rust. xoshiro256++ as the random source. Bit reservoir buffering. Tokio/Hyper wrapping HTTP. It returns a number every 0.5 to 0.8 nanoseconds. The maze is a 2D grid with walls and an endpoint. A request goes in. It returns a direction. No model. No weights. No GPU.

It cleared the maze in 892 steps. Under 300 ms.

Jev ran the same maze in 2,306 steps. It spent 2,295 of them in three cells at the (7,4) corner: right, down, right, down. The test timed out.

The tester gave Jev Manhattan distance as a heuristic. The input included last_move and visited_count for every neighboring cell. The instructions said: when a closer cell is blocked or visited too many times, pick the opening with the fewest visits.

Jev ignored it.

Right and down from the corner had shorter Manhattan distance. The fallback rule never fired. It stayed in a local optimum until timeout.

Two years of RLCD, two hours to reproduce

Diogo Almeida released Jev. He worked on ChatGPT/RLHF early. He spent two years in stealth mode developing RLCD: Reinforcement Learning for Calibrated Decisions.

Jev is a System One decision model. It does not write free text. It makes structured decisions, classifications, and routing calls. It returns type-safe results with calibrated probabilities. The pitch: 20 to 200 times faster than ordinary LLMs, 40 to 400 times cheaper, output tokens free. Built for software that needs to call a decision directly.

Harsha open-sourced Qwen-2.5-1B-RLCD. A small model based on Qwen-2.5-1B. It reproduced a stripped-down version of Jev’s core capability in a short time. TypeSafe AI’s Jev is closed source.

The two-hour reproduction optimized inference. It did not retrain a large model. Traditional LLM JSON generation runs token by token. It finishes one field before starting the next. Parallel constrained decoding evaluates all fields of the JSON structure in one forward pass. On an M4 MacBook, JSON workload inference improved 5 to 7 times. JSON schema compliance hit 100%.

Almeida spent two years on a training paradigm: calibrated probabilities, type-safe decisions, stable convergence across structured tasks. Harsha spent two hours on inference optimization: making a small model spit out compliant JSON under constrained decoding.

They are different races. The moat question remains. When core capabilities are easy to reproduce, model weights stop being the only barrier. Data loops, calibration quality, ecosystem integration, customer trust, and engineering reliability matter more.

Why the random strategy won

The random strategy’s maze win has math behind it.

Pólya’s random walk theorem: in a finite 2D connected grid, a simple random walk is recurrent. A drunk wandering a maze will eventually reach the endpoint with probability 1.

In a small 2D maze, with xoshiro256++ running at high QPS, the random strategy found the exit fast.

The theorem guarantees eventual arrival, not fast arrival. Make the maze larger and the step count explodes. A monkey at a typewriter can type Shakespeare. The publisher still waits forever.

The random strategy won on time. It lost on intelligence. It used brute-force search instead of planning. In a tiny maze, the recurrence of a 2D random walk did the work.

Jev’s failure is mismatch, not incompetence

Jev failed in the maze because it was used in the wrong place.

It is a fast single-step structured judge. Given a state and features, it picks among predefined options and returns a calibrated, type-safe decision. That is its job.

Detours, backtracking, multi-step planning, and counterintuitive choices when local heuristics conflict with global optimality are not its job.

A maze is the second kind of problem. Manhattan distance is a local heuristic. Most of the time, moving closer to the endpoint helps. In a maze with walls, you often move away first. Jev cannot understand “move away now so I can get closer later.” It picks the best-looking step. It gets stuck in the corner, choosing the shorter Manhattan distance until timeout.

That is the edge of System 1. A sprinter can play Go. He will not beat a Go player.

System 1 and System 2

Jev is System 1. Fast thinking: intuitive, automatic, parallel, low-energy.

Complex tasks need System 2. Slow thinking: rational, planning, serial, high-energy.

Mazes need System 2. Bin packing needs System 2. Scheduling needs System 2. Snake needs System 2. Autonomous driving needs System 2. Dropping a database, transaction operations, and audit scenarios need System 2 plus strict tool constraints and human oversight.

The workable architecture splits the job. System 2 plans. System 1 executes. Large models understand goals, break down tasks, and set strategy. Specialized decision models like Jev handle high-frequency, low-latency, structured single-step judgments. Tools handle deterministic operations. Memory and auditing handle traceability.

Caution list

Recommended: classification, routing, risk scoring, content moderation triage, simple decisions, and automated scenarios with high frequency, low latency, and tolerance for probabilistic errors.

Use caution or avoid: mazes, bin packing, scheduling, and other graph problems where local heuristics and global optimality disagree and that require detours, backtracking, or multi-step planning; sudden-death control like Snake, Tetris, and autonomous driving; irreversible high-cost decisions like dropping a database, transaction operations, and audit scenarios; complex tasks needing deep reasoning, global optimization, and interpretability.

Jev is a reflex arc for single-step decisions. It is not a navigator for multi-step planning.

Industry implications

AI is splitting into roles. General large models, specialized decision models, planning models, execution models, and tool-calling models each have a place. A group of specialized models working together will handle more problems than one large model.

Core capabilities commoditize fast. Once a capability works, open-source developers can reproduce it with small models, constrained decoding, and fine-tuning. The lead window shrinks. Data, calibration, reliability, ecosystem, and customer trust form the moat. Model weights alone do not.

Training paradigms and inference engineering are different. Two years in stealth explored training methods. Two hours reproduced inference optimization. Both matter. Their value differs. Founders should ask: are you exploring a new paradigm, or optimizing engineering?

Model choice depends on task fit. Do not worship two years in stealth. Do not worship a two-hour reproduction. Ask: does this model fit my task? Does my task need System 1 or System 2? Single-step or multi-step? Probability or determinism? Speed or correctness?

Hybrid systems work best. System 2 plans. System 1 decides. Tools execute. Memory audits.

Logs

The random-number server logs keep scrolling. 200 OK. 200 OK. 200 OK. Each response under a millisecond.

Jev’s log stops at (7,4). Right, down, right, down.

The test script stays. Next time the maze gets bigger, the random strategy will stall too. Placement decides the outcome.

how tosocial mediagadgetsthought leadersappsmobiletech newsfact or fictionfuture

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin