The Last Generation Paid Per Thought.
The DX-M2 Era
You are not renting intelligence anymore. You own it.
Ultra-low Power
~
0
W
World’s First 2nm
0
nm
Frontier-class models. Off the grid
~
0
B
NPU Performance
0
TOPS
Memory Bandwidth
0
GB/s
Faster than You Read (Single-Batch)
~
0
TPS
*Specifications are subject to change without notice during development.
Chip Line-up
How Big Should It Think?
Others count operations. We count how much mind fits on the die.
DX-M2FcBGA 16x16
External Memory
Up to 96GB
(LPDDR5X 24GB x 4EA)
Power Consumption
5W (Only SoC)
5W (Only SoC)
Max Bandwidth
8-Channel (153.6 GB/s)
8-Channel (153.6 GB/s)
Form Factor
M.2 Module or PCIe Card
M.2 Module or PCIe Card
DX-M2MFcMCM 19x19 (Tentative)
Internal Memory
Up to 24GB
(LPDDR5X 24GB x 1EA)
Power Consumption
5W (SoC 3W + DRAM 2W)
5W (SoC 3W + DRAM 2W)
Max Bandwidth
4-Channel (76.8 GB/s)
4-Channel (76.8 GB/s)
Form Factor
M.2 Module
M.2 Module
DX-M2M ProFcMCM 29x21 (Tentative)
Internal Memory
Up to 48GB
(LPDDR5X 24GB x 2EA)
Power Consumption
8.5W (SoC 4.5W + DRAM 4W)
8.5W (SoC 4.5W + DRAM 4W)
Max Bandwidth
8-Channel (153.6 GB/s)
8-Channel (153.6 GB/s)
Form Factor
M.2 Module
M.2 Module
Every Deployment has a Breakeven Point
Ours Arrives Sooner than You Think.
The cloud bills you for thinking. The DX-M2 bills you once.
The Cloud AI Breakeven Point
Cost Breakeven, per Device
The Cloud AI Breakeven Point
Dependency Trend by Model Size
Data Center Reliance (%)
Generative AI Model Size (Billion Parameters, 20B ~ 100B)
Shifting Workloads
Physical AI Takes the Edge
Deployment of Generative at the Edge
Cost Breakeven, per Device
Cumulative Inference Cost Over Time
Cumulative Cost ($)
Months since Purchase
Total Cost at 36 Months
Same device, same 3-year window
Total Savings: $315 Over 36 Months
Cloud Bottlenecks
Infrastructure expansion is constrained by power grid volatility and skyrocketing cloud costs.
Native Edge Demand
The market demands native LLM execution on high-performance edge devices.
The DX-M2 Solution
DX-M2 resolves these constraints with architectural innovations for next-gen physical AI.
DXNN SDK for DX-M2
Your Code Already Runs Here.
Nothing to port. Nothing to rewrite.
The compiler does the work you were planning to do.
The compiler does the work you were planning to do.
01
Connect
Bring Your Own Stack
Seamlessly run your existing frameworks and APIs on DX-M2 with zero code modification.
- OpenAI-compatible API
- PyTorch
- llama.cpp
- Ollama
- vLLM
- ExecuTorch
02
Optimize
Tuned to the Metal
Unlock peak hardware performance with an M2-native SDK that extracts every TOPS from silicon.
- DX-LLM: M2-native Runtime Framework
- Model / Memory / Compute Optimization
- NPU / DSP Compute Library
03
Unify
One Layer to Run It All
Consolidate standard frameworks and DX-LLM workloads under one unified middleware layer.
- M2 Middleware · Driver HAL
- Unified Compiler & Custom Kernel Build
Universal Model Ecosystem
DX-M2 Optimized Model Zoo.
Every model is pre-optimized for the DX-M2 NPU and served through DX-LLM.
These are the models we measured ourselves — quantization variants benchmarked
against the bf16 baseline, on-device, with zero token fees.
These are the models we measured ourselves — quantization variants benchmarked
against the bf16 baseline, on-device, with zero token fees.
Featured Models
Image-Text-to-Text
Alibaba
Qwen3-VL-8B-Instruct
Accuracy (PPL)
GPU Baseline
75.25
BF16
NPU Quantized
75.53
+0.28
W8A8
Text Generation
Alibaba
Qwen3-4B
Accuracy (PPL)
GPU Baseline
65.97
BF16
NPU Quantized
66.14
+0.16
W8A8
Explore All Models
Model
Task
Input
Output
Result
Qwen3-4BAlibaba
Text Generation
Text
Text
Deploy
Qwen3-VL-8B-InstructAlibaba
Image-Text-to-Text
Text
Image
Video
Text
Deploy