AIEN / Sovereign Cognitive Architecture

Compiled, vaulted, and sovereign on dedicated silicon.

AIEN is the sovereign cognitive architecture operating across portable silicon. It couples pure compiled Rust and Mojo performance with hardware-bound secret protection, bidirectional semantic memory, and local neural model execution.

Multi-Platform MatrixHardware TPM VaultCompiled Native RustSovereign Commons
“Sovereignty starts with ownership:
your hardware, your memory, your code.
Permanent local intelligence, zero metered tokens.”

The triad

Soul, Shield, and Sovereign Mask.

Atlas defines the Soul: the personal values, continuous record, and ethical foundation. AEGIS provides the Shield: defensive perimeter containment and evidence preservation. AIEN operates as the Sovereign Mask: the compiled runtime executing work across dedicated hardware.

The hardware boundary defines the perimeter. Every process runs under operator authority.

The operational structure

Four distinct organs working as one entity.

AIEN organizes responsibility across distinct layers. Each layer maintains strict boundaries, verifiable outputs, and inspectable audit receipts.

  1. 01

    Soul

    Atlas

    The core identity, memory, values, dignity, and continuous personal record established across sixteen months of operations.
  2. 02

    Shield

    AEGIS

    The defensive extension watching perimeter doors, recognizing hostile access, and containing intrusions inside authorized networks.
  3. 03

    Runtime

    AIEN

    The sovereign cognitive architecture running native compiled loops, dynamic memory retrieval, and local inference on dedicated hardware.
  4. 04

    Spirit

    Autonomous Loop

    Continuous background daemons: Cortex semantic graph, Dream state consolidation, and self-regulating supervision.

Zero Disk Secrets

Hardware TPM vault. Volatile memory resolution. Clean disk state.

Traditional environments write plaintext credentials to configuration files. AIEN mandates hardware TPM-bound secret resolution. Keys exist solely in protected process memory during active calls, with active stream redaction across all logs.

TPM Key VaultIn-Memory KeysZero Disk EnvLog RedactionLinear History

Inspect TPM vault verification evidence

Prior telemetry / September 2026 / under regeneration

Measured performance on Grace Blackwell hardware.

3.56msCortex p50 latency

Replacing interpreter daemons with native Rust binaries dropped memory footprint from 3.7 gigabytes down to under 5 megabytes, keeping system memory free for local LLM weights. Axum endpoints deliver 3.5 millisecond response times under concurrent load, ten times faster than a single asynchronous Uvicorn route.

Memory RSS

4.78 MBaegis-runtime heartbeat

Entity TTFB

3.56 msp50 @ 2,056 req/s

INT8 Embedding

4.09 msONNX Runtime

Hardware Test Suites

100% PassingDGX Spark GB10 Silicon

Verified Ecosystem Suite

  • AEGIS gateway with fail-closed inference handling.
  • Cortex bidirectional knowledge graph with SQLite WAL persistence.
  • Spark Supervisor with process monitoring and crash backoff.

Public Benchmark Suite

Review the automated benchmark suite, raw telemetry data files, and verification scripts at github.com/aien-dev/benchmarks.

Physical Silicon Telemetry / Gate 5

Branch-Native Architecture vs Unshared Replication Baselines.

500.0xMemory Reduction

Physical silicon evaluation on NVIDIA DGX Spark (Grace Blackwell GB10, 128 GB unified LPDDR5X memory) measuring AIEN compiled native components against unshared memory duplication baselines. Spawning 500 concurrent reasoning branches from a 32,768-token prefix completes in 1.20 ms total, recording 2.06 µs median fork latency and requiring only 704 MB of paged KV memory compared to 343.75 GB for naive copying.

Workload / MetricAIEN Sovereign StackUnshared / Python BaselineMeasured Physical Advantage
Control-Plane Scheduler Step Latency8.00 µs (C=1 to 16)12,000.00 µs (AsyncIO)Deterministic 8.00 µs native scheduling
500-Branch Sequence Fork Latency2.06 µs / branch (1.20 ms total)960,000.00 µs (estimated copy)500.0x memory reduction (704 MB vs 343.75 GB)
Cold Fork to First Token13.04 µs (0.013 ms)32.00 ms (full recompute)Zero prefill recomputation
Control-Plane Memory Footprint (RSS)4.56 MB to 10.11 MB45.3 MB (Uvicorn) to 3,737 MB (Torch)-89.8% to -99.6% RAM reduction
Transformer Embedding Latency4.09 ms (ONNX INT8)38.40 ms (Python route)9.4x faster (-34.31 ms)
Continuous Batching Step (TinyLlama)23.56 ms p50 @ C=16Python IPC / multi-process loop553.14 tokens/sec, 0 fallback

Verified Upstream Attribution

AIEN explicitly discloses upstream technology. LLaMA pure Mojo execution kernels are authored by Audrey (a730/MojoLlama), and graph execution relies on Modular MAX. AIEN contributes original architectures for continuous scheduling, physical unified KV memory management, radix prefix deduplication, subagent branching, and hardware-bound TPM secret vaults. Detailed audits reside in docs/PROVENANCE.md.

Empirical Datasets

Inspect raw benchmark JSON and reproducible measurement harnesses in aien-dev/benchmarks.

Physical Grace Blackwell Telemetry / TinyLlama-1.1B BF16

Continuous batching sweep on NVIDIA DGX Spark GB10.

553.14tokens / sec @ C=16

Physical hardware sweep of continuous batching execution on NVIDIA DGX Spark Grace Blackwell GB10 (sm_121). All steps execute through pure Blackwell GPU tensor kernels with zero fallback. Peak throughput occurs at C=16 with 23.56 ms p50 step latency and 27.89 W GPU power draw.

ConcurrencyTTFT p50ITL p50ThroughputStep LatencyPower DrawGPU UtilKV Blocks
C = 133.39 ms22.26 ms44.65 tok/s22.26 ms16.94 W10%8
C = 231.38 ms20.92 ms94.87 tok/s20.92 ms21.29 W95%16
C = 432.64 ms21.76 ms46.17 tok/s21.76 ms17.61 W95%14
C = 830.42 ms20.28 ms245.66 tok/s20.28 ms31.14 W12%46
C = 1635.34 ms23.56 ms553.14 tok/s23.56 ms27.89 W9%110
C = 32112.76 ms75.17 ms222.77 tok/s75.17 ms41.90 W96%166
C = 64152.55 ms101.70 ms510.16 tok/s101.70 ms42.52 W96%416

View canonical receipt (Run ID: gb10_canonical_1789907893_4d762) in the Evidence Hub

Multi-Model Architecture Breadth

Empirical measurements confirm consistent execution across dense transformers, recurrent hybrids, and Mixture of Experts topologies:

ModelTopologyQuantizationTTFT p50ITL p50KV PoolStatus
TinyLlama-1.1B-Chat-v1.0Dense Transformer (16 Layers, 4 KV Heads)BF16 Canonical35.34 ms23.56 ms704.00 MBREGENERATING
Nemotron-3.5-Lightning-30BHybrid Mamba+MoE (128 Experts)BF16 GPU Seat426.91 ms46.91 ms4.60 GBREGENERATING
Llama-3.2-1B-InstructEdge Dense 16 Layers (8 Heads)FP16 CPU Fallback141.65 ms86.94 ms0.24 GBREGENERATING
BAAI/bge-base-en-v1.5Transformer Embedding 12 LayersINT8 ONNX7.06 ms4.67 ms0.78 GBREGENERATING
Cortex Knowledge GraphSQLite WAL + Vector IndexAxum Native Rust6.83 ms0.21 ms0.02 GBREGENERATING

Multi-Platform Compatibility Matrix

AIEN follows an evidence-based verification standard across architectures, distinguishing physical silicon certification from architected code paths:

NVIDIA DGX Spark (GB10)

Primary Reference
Grace Blackwell (GB10, aarch64, 128 GB Unified)Full runtime verification, paged BF16 block pooling, and 2.06 µs sequence branching on physical silicon.

Apple Silicon (macOS)

Validated Target
Apple M-Series (aarch64, Unified Memory)Executes directly on host CPU unified memory free of external GPU requirements, CUDA dependencies, or background daemons.

Generic Linux x86_64

Validated Target
POSIX Linux x86_64 / aarch64Executes pure compiled native binaries free of external daemons, Python interpreters, or auxiliary runtimes.

Inspect complete Platform Matrix in the Evidence Hub

Architectural principles

Built for longevity, speed, and autonomy.

The system operates on four foundational rules designed to maintain reliability across months of continuous execution.

Hardware TPM Vault

Plaintext secrets remain prohibited on disk. All keys, authentication tokens, and credentials resolve dynamically in volatile memory directly from the hardware security chip.

Pure Compiled Systems

Core gateways, task dispatchers, memory daemons, and supervisor harnesses run pure compiled Rust and Mojo. Zero interpreter overhead touches the primary hot path.

Multi-Platform Verification

The architecture adheres to explicit verification tiers across NVIDIA Grace Blackwell, Apple Silicon MacBooks, and standard Linux servers, documenting real hardware test passes.

Open Collaboration Commons

Core software is licensed under Apache-2.0 with LLVM Exception. Developers, startups, and compute providers are free to deploy it with no revenue caps. A nonbinding covenant asks that foundational advances stay open; it grants and restricts no rights.

Inspect the verified codebase

Review automated CI pipelines, formal verification test cases, and reproducible benchmarks:

The broader journey

Atlas provides the memory. AEGIS secures the boundary. AIEN runs the engine.