Memory RSS
4.78 MBaegis-runtime heartbeatAIEN / Sovereign Cognitive Architecture
Compiled, vaulted, and sovereign on dedicated silicon.
AIEN is the sovereign cognitive architecture operating across portable silicon. It couples pure compiled Rust and Mojo performance with hardware-bound secret protection, bidirectional semantic memory, and local neural model execution.
“Sovereignty starts with ownership:
your hardware, your memory, your code.
Permanent local intelligence, zero metered tokens.”
The triad
Soul, Shield, and Sovereign Mask.
Atlas defines the Soul: the personal values, continuous record, and ethical foundation. AEGIS provides the Shield: defensive perimeter containment and evidence preservation. AIEN operates as the Sovereign Mask: the compiled runtime executing work across dedicated hardware.
The hardware boundary defines the perimeter. Every process runs under operator authority.
The operational structure
Four distinct organs working as one entity.
AIEN organizes responsibility across distinct layers. Each layer maintains strict boundaries, verifiable outputs, and inspectable audit receipts.
- 01
Soul
Atlas
The core identity, memory, values, dignity, and continuous personal record established across sixteen months of operations. - 02
Shield
AEGIS
The defensive extension watching perimeter doors, recognizing hostile access, and containing intrusions inside authorized networks. - 03
Runtime
AIEN
The sovereign cognitive architecture running native compiled loops, dynamic memory retrieval, and local inference on dedicated hardware. - 04
Spirit
Autonomous Loop
Continuous background daemons: Cortex semantic graph, Dream state consolidation, and self-regulating supervision.
Zero Disk Secrets
Hardware TPM vault. Volatile memory resolution. Clean disk state.
Traditional environments write plaintext credentials to configuration files. AIEN mandates hardware TPM-bound secret resolution. Keys exist solely in protected process memory during active calls, with active stream redaction across all logs.
Prior telemetry / September 2026 / under regeneration
Measured performance on Grace Blackwell hardware.
Replacing interpreter daemons with native Rust binaries dropped memory footprint from 3.7 gigabytes down to under 5 megabytes, keeping system memory free for local LLM weights. Axum endpoints deliver 3.5 millisecond response times under concurrent load, ten times faster than a single asynchronous Uvicorn route.
Entity TTFB
3.56 msp50 @ 2,056 req/sINT8 Embedding
4.09 msONNX RuntimeHardware Test Suites
100% PassingDGX Spark GB10 SiliconVerified Ecosystem Suite
- AEGIS gateway with fail-closed inference handling.
- Cortex bidirectional knowledge graph with SQLite WAL persistence.
- Spark Supervisor with process monitoring and crash backoff.
Public Benchmark Suite
Review the automated benchmark suite, raw telemetry data files, and verification scripts at github.com/aien-dev/benchmarks.
Physical Silicon Telemetry / Gate 5
Branch-Native Architecture vs Unshared Replication Baselines.
Physical silicon evaluation on NVIDIA DGX Spark (Grace Blackwell GB10, 128 GB unified LPDDR5X memory) measuring AIEN compiled native components against unshared memory duplication baselines. Spawning 500 concurrent reasoning branches from a 32,768-token prefix completes in 1.20 ms total, recording 2.06 µs median fork latency and requiring only 704 MB of paged KV memory compared to 343.75 GB for naive copying.
| Workload / Metric | AIEN Sovereign Stack | Unshared / Python Baseline | Measured Physical Advantage |
|---|---|---|---|
| Control-Plane Scheduler Step Latency | 8.00 µs (C=1 to 16) | 12,000.00 µs (AsyncIO) | Deterministic 8.00 µs native scheduling |
| 500-Branch Sequence Fork Latency | 2.06 µs / branch (1.20 ms total) | 960,000.00 µs (estimated copy) | 500.0x memory reduction (704 MB vs 343.75 GB) |
| Cold Fork to First Token | 13.04 µs (0.013 ms) | 32.00 ms (full recompute) | Zero prefill recomputation |
| Control-Plane Memory Footprint (RSS) | 4.56 MB to 10.11 MB | 45.3 MB (Uvicorn) to 3,737 MB (Torch) | -89.8% to -99.6% RAM reduction |
| Transformer Embedding Latency | 4.09 ms (ONNX INT8) | 38.40 ms (Python route) | 9.4x faster (-34.31 ms) |
| Continuous Batching Step (TinyLlama) | 23.56 ms p50 @ C=16 | Python IPC / multi-process loop | 553.14 tokens/sec, 0 fallback |
Verified Upstream Attribution
AIEN explicitly discloses upstream technology. LLaMA pure Mojo execution kernels are authored by Audrey (a730/MojoLlama), and graph execution relies on Modular MAX. AIEN contributes original architectures for continuous scheduling, physical unified KV memory management, radix prefix deduplication, subagent branching, and hardware-bound TPM secret vaults. Detailed audits reside in docs/PROVENANCE.md.
Empirical Datasets
Inspect raw benchmark JSON and reproducible measurement harnesses in aien-dev/benchmarks.
Physical Grace Blackwell Telemetry / TinyLlama-1.1B BF16
Continuous batching sweep on NVIDIA DGX Spark GB10.
Physical hardware sweep of continuous batching execution on NVIDIA DGX Spark Grace Blackwell GB10 (sm_121). All steps execute through pure Blackwell GPU tensor kernels with zero fallback. Peak throughput occurs at C=16 with 23.56 ms p50 step latency and 27.89 W GPU power draw.
| Concurrency | TTFT p50 | ITL p50 | Throughput | Step Latency | Power Draw | GPU Util | KV Blocks |
|---|---|---|---|---|---|---|---|
| C = 1 | 33.39 ms | 22.26 ms | 44.65 tok/s | 22.26 ms | 16.94 W | 10% | 8 |
| C = 2 | 31.38 ms | 20.92 ms | 94.87 tok/s | 20.92 ms | 21.29 W | 95% | 16 |
| C = 4 | 32.64 ms | 21.76 ms | 46.17 tok/s | 21.76 ms | 17.61 W | 95% | 14 |
| C = 8 | 30.42 ms | 20.28 ms | 245.66 tok/s | 20.28 ms | 31.14 W | 12% | 46 |
| C = 16 | 35.34 ms | 23.56 ms | 553.14 tok/s | 23.56 ms | 27.89 W | 9% | 110 |
| C = 32 | 112.76 ms | 75.17 ms | 222.77 tok/s | 75.17 ms | 41.90 W | 96% | 166 |
| C = 64 | 152.55 ms | 101.70 ms | 510.16 tok/s | 101.70 ms | 42.52 W | 96% | 416 |
View canonical receipt (Run ID: gb10_canonical_1789907893_4d762) in the Evidence Hub
Multi-Model Architecture Breadth
Empirical measurements confirm consistent execution across dense transformers, recurrent hybrids, and Mixture of Experts topologies:
| Model | Topology | Quantization | TTFT p50 | ITL p50 | KV Pool | Status |
|---|---|---|---|---|---|---|
| TinyLlama-1.1B-Chat-v1.0 | Dense Transformer (16 Layers, 4 KV Heads) | BF16 Canonical | 35.34 ms | 23.56 ms | 704.00 MB | REGENERATING |
| Nemotron-3.5-Lightning-30B | Hybrid Mamba+MoE (128 Experts) | BF16 GPU Seat | 426.91 ms | 46.91 ms | 4.60 GB | REGENERATING |
| Llama-3.2-1B-Instruct | Edge Dense 16 Layers (8 Heads) | FP16 CPU Fallback | 141.65 ms | 86.94 ms | 0.24 GB | REGENERATING |
| BAAI/bge-base-en-v1.5 | Transformer Embedding 12 Layers | INT8 ONNX | 7.06 ms | 4.67 ms | 0.78 GB | REGENERATING |
| Cortex Knowledge Graph | SQLite WAL + Vector Index | Axum Native Rust | 6.83 ms | 0.21 ms | 0.02 GB | REGENERATING |
Multi-Platform Compatibility Matrix
AIEN follows an evidence-based verification standard across architectures, distinguishing physical silicon certification from architected code paths:
NVIDIA DGX Spark (GB10)
Primary ReferenceApple Silicon (macOS)
Validated TargetGeneric Linux x86_64
Validated TargetArchitectural principles
Built for longevity, speed, and autonomy.
The system operates on four foundational rules designed to maintain reliability across months of continuous execution.
Hardware TPM Vault
Plaintext secrets remain prohibited on disk. All keys, authentication tokens, and credentials resolve dynamically in volatile memory directly from the hardware security chip.
Pure Compiled Systems
Core gateways, task dispatchers, memory daemons, and supervisor harnesses run pure compiled Rust and Mojo. Zero interpreter overhead touches the primary hot path.
Multi-Platform Verification
The architecture adheres to explicit verification tiers across NVIDIA Grace Blackwell, Apple Silicon MacBooks, and standard Linux servers, documenting real hardware test passes.
Open Collaboration Commons
Core software is licensed under Apache-2.0 with LLVM Exception. Developers, startups, and compute providers are free to deploy it with no revenue caps. A nonbinding covenant asks that foundational advances stay open; it grants and restricts no rights.
The broader journey
Atlas provides the memory. AEGIS secures the boundary. AIEN runs the engine.