Journal · Article 01

Why I’m Rethinking the Computer From the Language Down

The research began with practical questions about how work is represented and executed. Following those questions has expanded the investigation across language, execution, runtime, information representation, and the longer-term possibility of purpose-built computer architecture.

The question I started with

I did not begin with a finished execution architecture or a claim that I had found a faster computer. The work began on July 7, 2026, with a narrower question about organizing work in software: could useful information about a task—what kind of work it was, who or what owned it, where it should run, and how it related to other work—remain available long enough to influence execution?

The first prototypes were hardware-aware routing experiments. They discovered the local host, organized CPU work into lanes, and compared sequential, category-bridge, and microtask paths. The next versions added label-directed CPU and DirectML routing, then production-shaped tensor pipelines and real local Qwen workflows. At that stage the investigation was still application-level orchestration. The labels and routing rules lived around the work; they were not yet the language-and-execution contract that I now call NRL and NSE.

One result was especially useful because it was negative. In a fairness-controlled local Qwen test, three repeated comparisons put the category-owned path at roughly 1.00027× the ordinary concurrent workflow—effectively no meaningful advantage in that case. A later RAG-shaped local workflow produced a bounded improvement under its recorded conditions, but that did not rescue the original idea as a general claim. It showed that workload shape mattered. Concurrency by itself was not the answer I was looking for.

The problem moved below the application

Those experiments changed the question. Application code can label a task, choose a lane, or route a request, but that does not necessarily preserve the identity, relationships, changed state, persistence, priority, or intended destination of the work once it passes into lower layers. If those facts disappear before execution, a runtime has less information with which to decide what actually needs to happen.

I therefore moved from asking how to coordinate ordinary tasks to asking what information about declared work could survive into execution. That is a different problem. It touches the way a program describes work, the artifacts a compiler produces, the contract a runtime follows, and the evidence needed to show that changed inputs lead to the correct affected work without silently changing the program’s meaning.

This shift is the beginning of the current semantic-execution line. It does not imply that established computers are wrong or that every program should use a new model. It defines a testable research boundary: preserve more declared structure, make the resulting artifacts inspectable, and measure whether selective execution remains correct and useful under explicit conditions.

Why a language layer emerged

Once the work moved below application-level routing, I needed a language-facing surface that could state application work explicitly enough to test the execution question. That surface is Nioret Semantic Language, or NRL. Current source files use the .nrl extension. Historical names and extensions belong to precursor work and are not the current terminology.

NRL is not a decorative syntax over an otherwise finished system. It is the part of the research where typed application work is expressed. NSE is the related execution-model and compiler/runtime investigation. Keeping those names separate matters: the language describes work; the execution research investigates how declared relationships can remain present in deterministic, inspectable artifacts.

The V1 foundation dated September 28, 2026, is implemented rather than merely proposed. Its supported subset can check, inspect, compile, build, and run through native C++20. It covers typed modules, functions and semantic state, structures, local binding and mutation, field and index access, aggregates, assignments, common control flow, strings, option/list/map forms, file and persistent-text utilities, and local HTTP, form, and HTML work.

Two .nrl TaskBoard applications exercise that foundation: an interactive native application and a local HTTP/browser application with persistence. The recorded package passed 5 of 5 NRL checks and 28 of 28 NSE checks. That is a concrete language foundation, but it is still bounded. Imports and package resolution, generics, traits, Result, pattern matching, closures, asynchronous work, concurrency, raw systems facilities, a C ABI, WASM output, and browser-native output are not established.

Language alone was not the end of the problem

Creating a language could not answer the execution question by itself. NRL describes application work. Nioret Semantic Execution carries declared relationships into deterministic, inspectable execution artifacts and investigates whether an update can execute the work affected by changed inputs rather than treating every update as full recomputation.

The current implementation makes that distinction concrete. It can accept a bounded declared graph, produce native CPU artifacts, and select affected work for the graph forms it supports. CPU and CUDA are downstream physical backends; they are not the definition of the semantic contract. General NRL control flow does not yet become arbitrary NSE selective execution, so the connection between the language and the execution model is real but incomplete.

I am keeping the mechanism-level details out of this article. The public-safe account can explain what enters the system, what artifacts are produced, what correctness was tested, and where the model fails. It does not need to disclose how the compiler constructs its internal representations, transformations, dependencies, schedules, or generated execution paths.

What exists today

The current NSE software authority is the September 28, 2026 Gate 6 diligence package. Gate 0 modularized the earlier compiler and reproduced the accepted semantic map and native C++ output behavior. Gate 1 established deterministic CPU artifacts and structured diagnostics, then checked changed-input correctness across five graph shapes and all 35 nonzero change masks used by the suite at a 1e-10 tolerance.

The 1.0.0a4 alpha SDK exposes Python and command-line flows for graph and source inputs and produces inspectable CPU and CUDA output bundles. A separate C++ consumer compiled and ran the generated CPU output. That is demonstrated alpha tooling, not a production platform.

A controlled benchmark package also exists. It covers eight synthetic, auditable elementwise workload families and compares staged, fused, manually selective, and NSE-selective paths. Correctness was checked before timing, followed by 31 paired randomized trials per method at one and five threads in one recorded CPU environment. The quantitative result remains under benchmark and IP publication review, so I am not reproducing it here. The evidence does not establish production-application performance, universal speed, superiority to mature optimizing compilers, or independent reproduction.

The implementation can emit deterministic CUDA source, and its source-generation checks passed. The recorded Gate 4 environment had no CUDA compiler or device, however, so there is no physical CUDA build, execution, correctness, or performance result. Device qualification remains in development.

The results that changed my direction

The July fairness result ruled out a simple story in which category-owned concurrency was automatically better than ordinary concurrency. The September planner work produced a second important correction. An early approach materialized every changed-root combination. At 16 roots it produced 65,535 plans and a generated header larger than 58 million bytes. At 20 roots it reached 1,048,575 plans, approximately 630,424 KiB of peak memory, and about 5.02 seconds in the recorded probe. That approach did not scale.

The replacement sparse path stopped materializing every combination for sparse changes. Recorded shared and independent 4,096-root cases created one and 4,096 reusable fragments respectively, with no exhaustive plans, and a generated 1,024-root C++ header passed syntax checking. That corrected the specific scaling failure, but it did not create a universal policy. A dense-change probe showed that selective fragment execution can be worse than simply running the full path. Choosing adaptively between selective and full execution remains open work.

These failures narrowed the research in a useful way. I no longer treat routing, concurrency, selective execution, or one favorable workload as interchangeable proof. Each layer needs its own correctness boundary and its own comparison. The NRL foundation is therefore described as a supported application subset, the NSE SDK as alpha, the benchmark as synthetic and single-environment, and the CUDA path as emitted source rather than a qualified device backend.

Why this eventually reaches hardware

The progression from language to execution model to runtime naturally raises a hardware question. If declared relationships remain useful through software evaluation, it becomes reasonable to ask whether conventional processors expose the right primitives for that model or whether a purpose-built architecture could eventually carry the contract more directly.

That is a research direction, not a current result. Every demonstrated NSE path in the reviewed package runs on conventional hardware. No native NSE processor is implemented. The current CPU path is the physically compiled and executed backend; CUDA evidence stops at deterministic source emission.

The long-term target is NSE-oriented compute that can be evaluated against larger AI workloads, simulation, and scientific or high-density computation. Those domains are intended test areas, not promises that a future system will outperform established hardware. The software evidence has to broaden, the language-to-execution integration has to mature, and physical backends have to be qualified before custom hardware can be more than a research target.

Storage and information representation

The same line of inquiry also raises questions about information representation: what must be stored explicitly, what can be reconstructed, which external state is required, and what computation is necessary to recover the original information. At present, the reviewed evidence does not support a general NSE storage-density result, so I am not making one.

Nioret has a separate bounded animation-representation experiment that exactly reconstructed its recorded finite corpus and held-out set. That result is not evidence of a general storage architecture and should not be generalized into one. Any future storage or information-density claim will need precise accounting for physical bytes, logical information, metadata, dictionaries or models, external state, recoverability, reconstruction cost, content or entropy dependence, and failure cases. The underlying representation mechanisms remain outside this article.

What I am not publishing

Some parts of this research remain proprietary while patent and trade-secret review continues. I am not publishing compiler transformations, execution internals, internal representations, scheduling or dependency mechanisms, storage mechanisms, or processor mechanisms here. I am also not presenting a filing number or legal conclusion in this article.

This boundary is not meant to make the work immune from scrutiny. The goal is to expose versions, demonstrated behavior, methods, artifacts, hashes, failures, and limitations without also publishing the implementation recipe. A claim should be reviewable even when the mechanism that may carry patent or trade-secret value remains private.

How I expect the work to be evaluated

Future result records should identify the implementation version and commit, hardware and operating environment, compiler and build settings, workload, baseline, correctness checks, timing method, repetitions, per-run results, variance or dispersion, artifact hashes, known failures, and current limitations. Demonstrations should state exactly what was exercised and what the demonstration does not establish.

The present Gate 3 package follows part of that model: it records a specific implementation commit, one CPU environment, eight synthetic workload families, correctness checks, paired randomized trials, and a hashed result artifact. It is still one synthetic corpus on one machine, with several sub-millisecond cases and no independent reproduction. Those limits travel with the evidence.

Over time, evaluation should add broader workloads, physical backend qualification, additional machines, technical-record reconciliation, inspectable demonstrations, and independent reproduction where disclosure permits it. I do not claim that external verification has already happened.

Where the research is now

Today, NRL has an implemented V1 foundation and two working TaskBoard applications. NSE has a deterministic CPU artifact path, changed-input correctness evidence for its bounded graph model, a corrected sparse-planning path, an alpha SDK with an independently compiled CPU consumer, deterministic CUDA source emission, and a controlled synthetic benchmark package under publication review.

The incomplete parts are just as important. General NRL control flow does not yet derive arbitrary NSE selective execution. Dense-change policy is unfinished. The SDK remains alpha. CUDA has not been compiled or run on a device in the recovered evidence. Production workloads, broader machines, independent reproduction, and a public quantitative benchmark record remain outstanding. No native NSE processor exists, and there is no evidence-ready general storage claim.

The immediate work is therefore not to announce a finished replacement for conventional computing. It is to deepen NRL-to-NSE integration, improve the boundary between selective and full execution, qualify additional backends and workloads, and prepare disclosure-safe evidence records that other people can inspect. The longer-term architecture question remains open because the software work has made it more precise—not because it has already answered it.