Skip to content

Repository files navigation

NEXUS — AI-First Heterogeneous Accelerator

CI SystemVerilog Tests

NEXUS is an AI-first heterogeneous accelerator architecture designed from the ground up for LLM training/inference, reasoning models, Transformers, MoE, multimodal, diffusion, vision, speech, embedding/recommendation, scientific AI, and HPC tensor workloads.

It is not a traditional GPU. It is a hybrid: GPU-style parallelism + TPU-style matrix efficiency + NPU-style specialization + dataflow execution + a deep memory hierarchy + high-speed interconnect — co-designed with its ISA, compiler, runtime, and PyTorch backend.

Quickstart — everything runs, today

git clone https://github.com/devtantai-coder/nexus-ai-accelerator.git
cd nexus-ai-accelerator
bash verification/run_all.sh        # 22/22 PASS: vectors, T1/T2 models, 12 RTL testbenches

Requires only iverilog (Icarus Verilog ≥ 11) and python3 — no hardware, no EDA licenses. In ~30 seconds you get: 5 golden-vector generators, the analytic roofline model (T1), the cycle-approximate workload simulator with power model (T2), the compiler selftest, an end-to-end graph→NXT1-binary→execute loop, and 12 RTL testbenches — 4 of them bit-exact.

Simulate a 70B-class model end to end, with energy per token:

python3 simulator/t2/nexus_sim.py --workload decode_b1    # tok/s, W, J/token
python3 simulator/t2/nexus_sim.py --workload prefill_8k
python3 benchmarks/production_opt.py                      # config sweep incl. tok/s/W

What this is (honest status)

A pre-silicon architecture with executable evidence — roughly the first 15–20% of the road to a product (TRL 3): every claim traces to a running model or a passing testbench. No silicon exists yet. RTL modules are verified at proof scale (e.g. systolic array N=8 vs spec 128). See verification/README.md for the full evidence table and every bug the tests caught.

Layer Status
Spec + ISA complete, cross-checked (T1 caught 5 spec errors)
T1/T2 + power model all selftests PASS; J/token + tok/s/W per workload
RTL 10 modules, 12 testbenches, >9k golden vectors, 4 suites bit-exact
Compiler / runtime / NXT1 alpha prototypes, end-to-end loop runs
PyTorch backend proof-of-concept

Read this first

Every performance number in this repository is labeled [THEORETICAL], [ESTIMATED], [MEASURED], [HYPOTHESIS], [RESEARCH], [PRACTICAL], [UNKNOWN], or [INSUFFICIENT DATA]. Nothing is claimed until it is run.

Repository layout

nexus/
├── docs/          Architecture specification and design documents
├── rtl/           SystemVerilog RTL
│   ├── tensor/    Tensor Engine (128×128 systolic MMA, accumulator RF, formats)
│   ├── vector/    Vector Engine (1024 lanes, SFU, reductions)
│   ├── scalar/    Scalar RISC core (4-wide OoO, NEXUS ISA control)
│   ├── scheduler/ NCC task scheduler + Global Scheduler + barrier network
│   ├── cache/     L1/L2/SLC controllers, coherence directory
│   ├── sram/      Banked local SRAM wrappers, bank arbitration
│   ├── memory/    LSU, coalescing, device MMU, prefetch, KV engine
│   ├── hbm/       HBM3E memory controllers, PHY wrappers, link ECC
│   ├── noc/       4×4 mesh routers, IO-die crossbar, QoS
│   ├── dma/       Async copy engines, collective DMA
│   ├── attention/ Fused attention pipeline (FA2 datapath, mask gen, online softmax)
│   ├── moe/       Router unit, top-k network, permutation crossbar
│   ├── quant/     DPE statistics, quantize/dequantize, compression engine
│   └── top/       Package / chiplet / IO-die / NCC top levels
├── isa/           NEXUS ISA v1 formal definition (encoding, async-group memory model)
├── compiler/      LLVM/MLIR-based compiler (nexus.graph/tensor/flow dialects), NKL
├── runtime/       NEXUS Runtime (streams, graphs, memory pools, collectives, profiler)
├── driver/        Kernel-mode driver (command rings, doorbells, device management)
├── pytorch/       torch.device("nexus") backend (eager op library + torch.compile path)
├── kernels/       Hand-tuned NKL kernel library (backstop & reference implementations)
├── simulator/     Architectural simulator (T1 analytic, T2 cycle-approx, T3 corr.)
├── fpga/          FPGA prototype (1 NCC + scheduler; software bring-up platform)
├── benchmarks/    Benchmark suite (§29/B of the spec) + roofline tooling
├── verification/  UVM testbenches, formal property sets, co-sim harness
└── docs/          This specification and future design notes

Design law

Effective Performance = Compute × Data Reuse × Memory Bandwidth
                      × Parallelism × Interconnect × Compiler

No compute block is added unless the memory hierarchy and interconnect can feed it (roofline-normative, spec §1). Priority order: memory efficiency → tensor utilization → data reuse → kernel fusion → interconnect → compiler → power → raw compute.

X1 headline targets (all [THEORETICAL]/[ESTIMATED] — see spec §2.2)

Package 4× N3E compute chiplets + N6 IO die + 8× HBM3E (2.5D bridge-class)
BF16 dense peak 1.78 PFLOPS (FP8 3.57 PF, FP4 7.13 PF)
HBM 256 GB @ 9.83 TB/s
On-package SRAM 240 MB
Scale-up NXL 1.8 TB/s per package
TDP 1200 W liquid / 900 W air

What actually runs today

Layer Artifact Evidence
Spec docs/NEXUS_Architecture_Specification_v1.0.md (30 sections) + isa/nexus_isa_v1.md (normative, v1.0 erratum 1) T1 cross-checks caught 5 spec errors
Analytic model (T1) simulator/t1/nexus_roofline.py 29/29 checks PASS
Executable sim (T2) simulator/t2/nexus_sim.py — runs LLM workloads (decode/prefill 70B-class) and NXT1 binaries selftest ALL PASS; within 25% of T1; prefill MFU 80% of roofline
RTL 10 modules: te_array (+AUTO_DRAIN), te_bf16_mac, te_fmt_cvt, te_fp8_mac, te_epilogue (SiLU/GELU), te_sparse_skip (2:4), td64_walker (pipelined), fp32_addsub, fp32_mul, exp_lut, softmax_online 11 testbenches PASS, >9k vectors, 4 suites bit-exact
Compiler compiler/nexus_graph.py (fusion/tiling/dataflow/roofline) + NXT1 emission (--emit) selftest ALL PASS
Runtime runtime/nxt1.py — shared task-binary codec round-trip tested
End-to-end loop graph → compile → NXT1 binary → T2 executes → performance report mlp/decode/attention.nxt all execute
RTL system test td64_walker drives te_array through a 2-group mma.group chain (tb_group_exec.sv) PASS, both C matrices exact

Energy-axis results (production decode, 70B, ctx 8K) [MODEL §27]

Config tok/s (b=1) J/token (b=1) tok/s/W (b=64)
BF16 baseline 57 8.77 5.66
FP8 weight 114 4.39 10.79
FP4 weight 223 2.24 19.57
FP4 + INT4 KV 227 2.20 19.77
FP4 + INT4 + 1.4× compress (prod max) 313 1.58 24.60

→ Production config cuts J/token 5.6× and lifts tok/s/W 4.3× vs BF16 — HBM traffic dominates decode energy (44–70% of package power), so quantization

  • compression attack the energy problem at exactly its physical root.

Performance upgrades landed (RTL v1.1, all regression-verified)

Upgrade Module Effect Evidence
Pipelined descriptor walk td64_walker 5–6 → 2 cycles/group (≥2.5× launch speedup on mma.group/MoE chains) tb_td64_walker, tb_group_exec PASS; hand-derived 2-cycle steady state
AUTO_DRAIN result streaming te_array drain overlaps wavefront tail: start→done −(N+2) cycles/GEMM (measured 42→34 @ N=8/K=16; −130/GEMM at N=128, up to ~26% on short-K decode/MoE tiles) tb_te_array_auto 3/3 PASS bit-exact, back-to-back reuse verified
Fused epilogue datapath te_epilogue bias + ReLU/SiLU/GELU real datapath (was passthrough stub) — removes per-layer activation pass over SRAM tb_te_epilogue 256×4 PASS (bias/relu bit-exact, silu/gelu within 2⁻⁹ LUT budget)
Real 2:4 sparse skip te_sparse_skip cluster-coded metadata, phase-accurate skip decisions (was trivial stub) tb_te_sparse_skip 6/6 PASS incl. 10/16 skip accounting
T2 drain-model sync nexus_sim.py per-tile drain 128 → 2 cycles [ESTIMATED from measured RTL], compress path refactored (no code duplication) T2 selftest ALL PASS, still within T1 tolerance
Power & energy model (§27 executable) nexus_power.py + T2 integration per-block watts (TE/VEC/SRAM/NoC/HBM/IO) from §27 budget, DVFS 4 points (f³ scaling), HBM 2-anchor PHY fit; J/token + tok/s/W in every report selftest: worst-case 1319 W ≤ TDP+10%, idle 144 W, decode 498 W < prefill 521 W (physical ordering)

Everything performance-related remains [THEORETICAL]/[ESTIMATED]: no silicon, no synthesis data, no [MEASURED] claims. RTL is v0 scale (arrays N=4/8 vs spec 128). See verification/README.md for the full evidence list and debug lessons.

About

Open AI accelerator architecture: roofline-disciplined spec, verified RTL (systolic tensor engine, AUTO_DRAIN, 2:4 sparsity), T1/T2 performance+power simulators, compiler & 22-test regression — all evidence-labeled

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages