- SpaDA: A Spatial Dataflow Architecture Programming Language
- Ripple: Asynchronous Programming for Spatial Dataflow Architectures
- TileLoom: Automatic Dataflow Planning for Tile-Based Languages on Spatial Dataflow Accelerators
- An MLIR Lowering Pipeline for Stencils at Wafer-Scale
- Neura: A Unified Framework for Hierarchical and Adaptive CGRAs
- Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
- CODO: An Automated Compiler for Comprehensive Dataflow Optimization
- Plasticine: A Reconfigurable Architecture for Parallel Patterns (ISCA 2017)
- Dato: A Task-Based Programming Model for Dataflow Accelerators
- TCP: A Tensor Contraction Processor for AI Workloads Industrial Product
- Cerebras Architecture Deep Dive: First Look Inside the Hardware/Software Co-Design for Deep Learning
- A software-defined tensor streaming multiprocessor for large-scale machine learning
- Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads
- In-Datacenter Performance Analysis of a Tensor Processing Unit
- MN-Core 2 White Paper