NEXUS is an AI-first heterogeneous accelerator architecture designed from the ground up for LLM training/inference, reasoning models, Transformers, MoE, multimodal, diffusion, vision, speech, embedding/recommendation, scientific AI, and HPC tensor workloads.
It is not a traditional GPU. It is a hybrid: GPU-style parallelism + TPU-style matrix efficiency + NPU-style specialization + dataflow execution + a deep memory hierarchy + high-speed interconnect — co-designed with its ISA, compiler, runtime, and PyTorch backend.
git clone https://github.com/devtantai-coder/nexus-ai-accelerator.git
cd nexus-ai-accelerator
bash verification/run_all.sh # 22/22 PASS: vectors, T1/T2 models, 12 RTL testbenchesRequires only iverilog (Icarus Verilog ≥ 11) and python3 — no hardware, no EDA
licenses. In ~30 seconds you get: 5 golden-vector generators, the analytic roofline
model (T1), the cycle-approximate workload simulator with power model (T2), the
compiler selftest, an end-to-end graph→NXT1-binary→execute loop, and 12 RTL
testbenches — 4 of them bit-exact.
Simulate a 70B-class model end to end, with energy per token:
python3 simulator/t2/nexus_sim.py --workload decode_b1 # tok/s, W, J/token
python3 simulator/t2/nexus_sim.py --workload prefill_8k
python3 benchmarks/production_opt.py # config sweep incl. tok/s/WA pre-silicon architecture with executable evidence — roughly the first 15–20% of
the road to a product (TRL 3): every claim traces to a running model or a passing
testbench. No silicon exists yet. RTL modules are verified at proof scale
(e.g. systolic array N=8 vs spec 128). See verification/README.md for the full
evidence table and every bug the tests caught.
| Layer | Status |
|---|---|
| Spec + ISA | complete, cross-checked (T1 caught 5 spec errors) |
| T1/T2 + power model | all selftests PASS; J/token + tok/s/W per workload |
| RTL | 10 modules, 12 testbenches, >9k golden vectors, 4 suites bit-exact |
| Compiler / runtime / NXT1 | alpha prototypes, end-to-end loop runs |
| PyTorch backend | proof-of-concept |
- NEXUS Architecture Specification v1.0 — the complete 30-section architecture specification (the primary deliverable).
Every performance number in this repository is labeled [THEORETICAL], [ESTIMATED],
[MEASURED], [HYPOTHESIS], [RESEARCH], [PRACTICAL], [UNKNOWN], or
[INSUFFICIENT DATA]. Nothing is claimed until it is run.
nexus/
├── docs/ Architecture specification and design documents
├── rtl/ SystemVerilog RTL
│ ├── tensor/ Tensor Engine (128×128 systolic MMA, accumulator RF, formats)
│ ├── vector/ Vector Engine (1024 lanes, SFU, reductions)
│ ├── scalar/ Scalar RISC core (4-wide OoO, NEXUS ISA control)
│ ├── scheduler/ NCC task scheduler + Global Scheduler + barrier network
│ ├── cache/ L1/L2/SLC controllers, coherence directory
│ ├── sram/ Banked local SRAM wrappers, bank arbitration
│ ├── memory/ LSU, coalescing, device MMU, prefetch, KV engine
│ ├── hbm/ HBM3E memory controllers, PHY wrappers, link ECC
│ ├── noc/ 4×4 mesh routers, IO-die crossbar, QoS
│ ├── dma/ Async copy engines, collective DMA
│ ├── attention/ Fused attention pipeline (FA2 datapath, mask gen, online softmax)
│ ├── moe/ Router unit, top-k network, permutation crossbar
│ ├── quant/ DPE statistics, quantize/dequantize, compression engine
│ └── top/ Package / chiplet / IO-die / NCC top levels
├── isa/ NEXUS ISA v1 formal definition (encoding, async-group memory model)
├── compiler/ LLVM/MLIR-based compiler (nexus.graph/tensor/flow dialects), NKL
├── runtime/ NEXUS Runtime (streams, graphs, memory pools, collectives, profiler)
├── driver/ Kernel-mode driver (command rings, doorbells, device management)
├── pytorch/ torch.device("nexus") backend (eager op library + torch.compile path)
├── kernels/ Hand-tuned NKL kernel library (backstop & reference implementations)
├── simulator/ Architectural simulator (T1 analytic, T2 cycle-approx, T3 corr.)
├── fpga/ FPGA prototype (1 NCC + scheduler; software bring-up platform)
├── benchmarks/ Benchmark suite (§29/B of the spec) + roofline tooling
├── verification/ UVM testbenches, formal property sets, co-sim harness
└── docs/ This specification and future design notes
Effective Performance = Compute × Data Reuse × Memory Bandwidth
× Parallelism × Interconnect × Compiler
No compute block is added unless the memory hierarchy and interconnect can feed it (roofline-normative, spec §1). Priority order: memory efficiency → tensor utilization → data reuse → kernel fusion → interconnect → compiler → power → raw compute.
| Package | 4× N3E compute chiplets + N6 IO die + 8× HBM3E (2.5D bridge-class) |
| BF16 dense peak | 1.78 PFLOPS (FP8 3.57 PF, FP4 7.13 PF) |
| HBM | 256 GB @ 9.83 TB/s |
| On-package SRAM | 240 MB |
| Scale-up | NXL 1.8 TB/s per package |
| TDP | 1200 W liquid / 900 W air |
| Layer | Artifact | Evidence |
|---|---|---|
| Spec | docs/NEXUS_Architecture_Specification_v1.0.md (30 sections) + isa/nexus_isa_v1.md (normative, v1.0 erratum 1) |
T1 cross-checks caught 5 spec errors |
| Analytic model (T1) | simulator/t1/nexus_roofline.py |
29/29 checks PASS |
| Executable sim (T2) | simulator/t2/nexus_sim.py — runs LLM workloads (decode/prefill 70B-class) and NXT1 binaries |
selftest ALL PASS; within 25% of T1; prefill MFU 80% of roofline |
| RTL | 10 modules: te_array (+AUTO_DRAIN), te_bf16_mac, te_fmt_cvt, te_fp8_mac, te_epilogue (SiLU/GELU), te_sparse_skip (2:4), td64_walker (pipelined), fp32_addsub, fp32_mul, exp_lut, softmax_online |
11 testbenches PASS, >9k vectors, 4 suites bit-exact |
| Compiler | compiler/nexus_graph.py (fusion/tiling/dataflow/roofline) + NXT1 emission (--emit) |
selftest ALL PASS |
| Runtime | runtime/nxt1.py — shared task-binary codec |
round-trip tested |
| End-to-end loop | graph → compile → NXT1 binary → T2 executes → performance report | mlp/decode/attention.nxt all execute |
| RTL system test | td64_walker drives te_array through a 2-group mma.group chain (tb_group_exec.sv) |
PASS, both C matrices exact |
| Config | tok/s (b=1) | J/token (b=1) | tok/s/W (b=64) |
|---|---|---|---|
| BF16 baseline | 57 | 8.77 | 5.66 |
| FP8 weight | 114 | 4.39 | 10.79 |
| FP4 weight | 223 | 2.24 | 19.57 |
| FP4 + INT4 KV | 227 | 2.20 | 19.77 |
| FP4 + INT4 + 1.4× compress (prod max) | 313 | 1.58 | 24.60 |
→ Production config cuts J/token 5.6× and lifts tok/s/W 4.3× vs BF16 — HBM traffic dominates decode energy (44–70% of package power), so quantization
- compression attack the energy problem at exactly its physical root.
| Upgrade | Module | Effect | Evidence |
|---|---|---|---|
| Pipelined descriptor walk | td64_walker |
5–6 → 2 cycles/group (≥2.5× launch speedup on mma.group/MoE chains) |
tb_td64_walker, tb_group_exec PASS; hand-derived 2-cycle steady state |
| AUTO_DRAIN result streaming | te_array |
drain overlaps wavefront tail: start→done −(N+2) cycles/GEMM (measured 42→34 @ N=8/K=16; −130/GEMM at N=128, up to ~26% on short-K decode/MoE tiles) | tb_te_array_auto 3/3 PASS bit-exact, back-to-back reuse verified |
| Fused epilogue datapath | te_epilogue |
bias + ReLU/SiLU/GELU real datapath (was passthrough stub) — removes per-layer activation pass over SRAM | tb_te_epilogue 256×4 PASS (bias/relu bit-exact, silu/gelu within 2⁻⁹ LUT budget) |
| Real 2:4 sparse skip | te_sparse_skip |
cluster-coded metadata, phase-accurate skip decisions (was trivial stub) | tb_te_sparse_skip 6/6 PASS incl. 10/16 skip accounting |
| T2 drain-model sync | nexus_sim.py |
per-tile drain 128 → 2 cycles [ESTIMATED from measured RTL], compress path refactored (no code duplication) | T2 selftest ALL PASS, still within T1 tolerance |
| Power & energy model (§27 executable) | nexus_power.py + T2 integration |
per-block watts (TE/VEC/SRAM/NoC/HBM/IO) from §27 budget, DVFS 4 points (f³ scaling), HBM 2-anchor PHY fit; J/token + tok/s/W in every report | selftest: worst-case 1319 W ≤ TDP+10%, idle 144 W, decode 498 W < prefill 521 W (physical ordering) |
Everything performance-related remains [THEORETICAL]/[ESTIMATED]: no silicon,
no synthesis data, no [MEASURED] claims. RTL is v0 scale (arrays N=4/8 vs
spec 128). See verification/README.md for the full evidence list and
debug lessons.