Hardware Sovereignty Telemetry
VRAM BOUNDED
GPU Accelerator NVIDIA GeForce GTX 1050 Ti
VRAM Allocation 3.82 GB / 4.00 GB (95.5%)
Analytical Left Brain ($T=0.1$) 0.84 GB VRAM
Creative Right Brain ($T=0.8$) 0.84 GB VRAM
Colossus & UKS Memory 2.14 GB VRAM
Thermal / Stability State 58Β°C β€’ Stable 60 FPS
Search... ⌘K Launch Studio

The Canonical B-Series Model Suite

Engineered under the Concentrated Intelligence Doctrine: maximizing semantic density per parameter to deliver professional-grade conversational and multimodal capabilities on sub-4GB consumer hardware.

Canonical Model Lineup and Builder Flow
Canonical B-Series Model Suite: B1 Hope (39M), B2 Insight (50M), B3 Apex (504M), B3 Ultra (3.2B MoE) MASTER MATRIX

Architectural Specifications by Model Tier

B1 Hope (39.2M)

0.23 GB VRAM

Ultra-low-latency device-local Small Language Model designed for instant command processing, text completion, and 40-second training epochs on GTX 1050 Ti.

Layers ($L$): 8Hidden Dim ($d_{model}$): 768
Attention Heads: 12 ($d_{head}=64$)Context Window ($T$): 4,096 tokens
FP16 Size: 78.4 MBINT8 Size: 39.2 MB
Inference Speed: ~55 tok/secTraining Epoch: ~40 seconds
Target: Consumer Laptops, Raspberry Pi, GTX 1050 Ti (4GB)

B2 Insight (50.1M)

0.35 GB VRAM

Intermediate cross-modal reasoning engine with Cross-Modal Cross-Attention (CMCA), aligning textual token streams with visual feature grids and acoustic spectrograms.

Layers ($L$): 10Hidden Dim ($d_{model}$): 832
Attention Heads: 13 (Latent)Context Window ($T$): 4,096 tokens
FP16 Size: 100.2 MBINT8 Size: 50.1 MB
Inference Speed: ~42 tok/secMultimodal: Text + Vision + Speech
Target: Smart Cameras, Edge Vision Systems, GTX 1050 Ti

B3 Apex (504.2M)

1.80 GB VRAM

Heavyweight edge foundation model featuring Multi-Head Latent Attention (MLA), FlashAttention-2, and deep Socratic conversational reasoning capabilities.

Layers ($L$): 24Hidden Dim ($d_{model}$): 3,072
Attention Heads: 24 ($d_{head}=128$)Context Window ($T$): 4,096 tokens
FP16 Size: 1.01 GBINT8 Size: 504.2 MB
Inference Speed: ~12 tok/sec (Pascal)Capabilities: Socratic Dialogue & Code
Target: GTX 1050 Ti (with INT8), RTX 3060/4060, Local Desktops

B3 Ultra MoE (3.2B)

3.80 GB (INT4)

Sovereign multimodal digital twin cognitive core. Features 8 experts with top-2 softmax gating, Manifold-Constrained Hyper-Connections, and 8,192 context window.

Total Params: 3.21 BillionActive Params: 852 Million
Layers: 32 (4 MoE Blocks)Experts: 8 Experts (Top-2 Routing)
INT4 GGUF Size: 1.61 GBContext Window ($T$): 8,192 tokens
Inference Speed: ~6 tok/sec (Quant)Primary Use: Lifelong Digital Twin
Target: GTX 1050 Ti (with INT4), Multi-GPU Edge, Enterprise Clusters

C1 Colossus & Forward (Commercial Flagship)

COMMERCIAL LICENSING

While the canonical B-Series suite (B1, B2, B3, B3 Ultra) remains 100% open-source under the permissive MIT License for individual researchers and the global developer community, C1 Colossus and all subsequent flagship models are commercially licensed for enterprise deployments, multi-datacenter clusters, and proprietary industrial integrations.

🏒 Enterprise Multi-Node Scale

Multi-GPU and multi-node cluster orchestration optimized for RTX 4090, H100, AMD Instinct, and Intel Gaudi hardware environments.

πŸ”’ Custom Knowledge Distillation

Bespoke teacher-student distillation pipelines tailored to proprietary corporate datasets, maintaining strict local data sovereignty.

πŸ“œ SLA & Commercial Warranties

Dedicated engineering support, commercial integration rights, indemnification, and direct migration pathways to ImpressionCore-S1.

Review Commercial Licensing & Dual-Tier Policy β†’ Community models supported by the voluntary $1 Sovereign AI Patron Pledge.

Live Model Architecture & VRAM Budget Calculator

Configure custom model parameters or select canonical presets to compute exact parameter counts, KV cache consumption, and GTX 1050 Ti VRAM feasibility in real time.

Interactive Model Parameter Simulator

Select a preset or customize parameters dynamically:

Transformer Layers ($L$) 8
Hidden Dimension ($d_{model}$) 768
Attention Heads ($n_{heads}$) 12
Context Window ($T$) 4096
Generated PyTorch Architecture Live Compilation

              
Total Parameters 39.2 Million
FP16 Weights Memory 0.08 GB
INT8 Quantized Memory 39 MB
INT4 GGUF Memory 20 MB
GTX 1050 Ti VRAM Budget (4.0 GB) 0.23 GB / 4.00 GB (6%)

πŸ’‘ Concentrated Intelligence: All calculations incorporate parameter weights, KV cache allocations, activation overhead, and gradient checkpointing savings.

Empirical Hardware Benchmark Matrix

Measured inference latency and memory utilization across consumer edge hardware to high-end workstations.

Hardware Platform VRAM / RAM B1 Hope (39M) B2 Insight (50M) B3 Apex (504M) B3 Ultra (3.2B MoE) VRAM Envelope Status
NVIDIA GTX 1050 Ti 4 GB VRAM ~55 tok/sec ~42 tok/sec ~12 tok/sec (INT8) ~6 tok/sec (INT4) ● 100% STABLE (<4GB)
NVIDIA RTX 3060 12 GB VRAM ~140 tok/sec ~110 tok/sec ~45 tok/sec ~28 tok/sec ● OPTIMAL HEADROOM
NVIDIA RTX 4090 24 GB VRAM ~280 tok/sec ~220 tok/sec ~115 tok/sec ~75 tok/sec ● MAXIMUM SPEED
Intel Core i5-4460 CPU 16 GB DDR3 ~45 tok/sec ~32 tok/sec ~6 tok/sec (GGUF) ~2.5 tok/sec (GGUF) ● CPU ZERO-GPU EXEC