The Banana Pi BPI-SM10 is a Pico-ITX board around the SpacemiT K3 SoC (CoM260 / COM3K316128…): 8× X100 general-purpose cores (RVA23, RVV 1.0, VLEN=256) plus 8× A100 AI cores (IME2, VLEN=1024, up to 60 TOPS), integrated IMG PowerVR BXM-4-64 GPU (Vulkan 1.3 / OpenCL 3.0). Ours runs Bianbu 4.0.6 with EESSI on dev.eessi.io/riscv.
Sponsored by Banana Pi — this board was provided by Banana Pi. Product page: BPI-SM10 (K3-CoM260) · docs.
X100 vs A100
K3 is two asymmetric RISC-V clusters. Both implement RVV 1.0 (A100 is not scalar-only); the AI cores trade general-purpose behaviour for a wider vector unit and IME2. Device-tree: spacemit,x100 (cpu@0–7) vs spacemit,a100 (cpu@8–15).
| X100 (cores 0–7) | A100 (cores 8–15) | |
|---|---|---|
| Role | General-purpose CPU | AI / matrix (IME2) |
| Clock (typ.) | up to 2.4 GHz | up to 2.0 GHz |
| Clusters | 0–1 (cluster_cpus 0–3, 4–7) |
2–3 (cluster_cpus 8–11, 12–15) |
| ISA (hart) | rv64imafdcvh + bitmanip / vector crypto |
rv64imafdcv (no h) + same vector ext. family |
| Hypervisor | yes (h) |
no |
| RVV | yes — VLEN=256 (vlenb=32, measured) |
yes — VLEN=1024 (vlenb=128, measured on cpu8) |
| Custom | — | IME2 · up to 60 TOPS (vendor matrix path) |
| Default Linux scheduling | yes — normal SMP | fenced — online in sysfs, but not in a normal task’s affinity mask |
Plain RVV on A100 is often slower than on X100 without vendor IME2 code — wide VLEN is there for the matrix path, not as a drop-in “faster X60.”
Getting onto A100 (/proc/set_ai_thread)
SpacemiT’s kernel keeps general threads on X100 even when A100s are idle. Stock taskset -c 8 fails with EINVAL until the thread is registered as an AI thread:
echo $$ > /proc/set_ai_thread # world-writable on Bianbu; unlocks cores 8–15 for this TID
taskset -c 8-15 … # now affinity works; process stays on A100 for its lifetime
On our board /proc/set_ai_thread is mode 0222 (write-only). After the write, Cpus_allowed_list becomes 8–15 (X100 is no longer available to that task).
VLEN pitfall: do not start on X100 (VLEN=256), let ld.so cache vector decisions, then migrate to A100 (VLEN=1024) — that pattern can SIGSEGV. Prefer set_ai_thread before exec of the real binary (same pattern as k3_taskset / SpacemiT’s AI runtimes).
For mixed MPI jobs (e.g. HPL), ranks on A100 each echo $TID > /proc/set_ai_thread then pin to cores 8–15 before exec. Equal rank counts are a poor fit for HPL — see heterogeneous balancing below.
Compute paths on this board
All four paths we care about are present — and the Custom / GPU story differs from K1 (X60 IME1 / BXE-2-32).
| Path | On SM10 / K3 |
|---|---|
| Scalar | X100 — rv64gc baseline under RVA23 |
| Vector | X100 RVV 1.0 · Zvl256b (same VLEN class as X60) |
| Custom | A100 cluster — IME2 matrix path · 60 TOPS (VLEN=1024) |
| GPU | BXM-4-64 — the SKU SpaceMiT’s proprietary OpenCL/Vulkan DDK actually targets (unlike K1’s BXE) |
Per-board diagrams for the other machines: VisionFive 2 · Orange Pi RV2 · Banana Pi F3. Homepage logo is the combined view.
Status
Board is on the desk and reachable over SSH with CVMFS + EESSI (EESSI_VERSION_OVERRIDE=2025.06-001, software tree riscv64/generic). eessi_archdetect reports SpacemiT X100. Default user threads only see cores 0–7; A100 work needs /proc/set_ai_thread.
HPL (OpenBLAS RVV)
See also the HPL app overview.
EESSI foss/2025b + HPL/2.3, FlexiBLAS → local OpenBLAS 0.3.34. Two RVV targets matter on K3:
RISCV64_ZVL256Bon X100 (VLEN=256), plus a small portability fix so stock ZVL256B wide-load +vgetLMUL splits (which assume VLMAX@256) do not corrupt tails when run at larger VLEN.RISCV64_ZVL1024Bon A100 (VLEN=1024) — new target: DGEMM16×8(+ S/C/Z/TRMM) from OpenBLAS’sgenerate_kernel.py,-march=…_zvl1024b. Patch: opensolvers/benchmarksOpenBLAS/OpenBLAS-0.3.34_add-riscv64-zvl1024b.patch.
OMP_NUM_THREADS=1 unless noted, NB=192, residual checks PASSED.
Single-cluster baselines (N=12000, 8 ranks, 2×4)
| Cluster | OpenBLAS target | GFLOP/s | Result |
|---|---|---|---|
| X100 only | ZVL256B | 52.0 | PASSED |
| A100 only | ZVL256B (VLEN-safe) | 15.0 | PASSED |
| A100 only | ZVL1024B | 35.9 | PASSED |
A100 ZVL1024B is ~2.4× the portable-ZVL256B A100 HPL rate (~1.45× slower than X100, vs ~3.5× before). Stock NETLIB on A100 was ~0.8 GFLOP/s @ N=4000.
A100 ZVL1024B at smaller N: 22.5 GF @ N=4000, 30.4 GF @ N=8000 (both PASSED).
Equal 16-rank mixed (8+8), same BLAS everywhere
Ranks 0–7 on X100, 8–15 on A100 (set_ai_thread), 4×4 grid:
| N | BLAS on A100 | GFLOP/s | Result |
|---|---|---|---|
| 4000 | ZVL256B | 16.2 | PASSED |
| 8000 | ZVL256B | 23.7 | PASSED |
| 12000 | ZVL256B | 27.5 | PASSED |
| 12000 | ZVL1024B (X100 still ZVL256B) | 57.5 | PASSED |
With ZVL256B on both clusters, equal ranks stall on slow A100 work. With ZVL1024B on A100, equal 8+8 becomes the best mixed config and beats X100-only.
Heterogeneous X100 + A100 (N=12000)
Per-rank FlexiBLAS: X100 → OPENBLAS_DUAL (ZVL256B), A100 → OPENBLAS_ZVL1024B. Wrap: HPL_NX100, HPL_A100_THREADS, set_ai_thread + taskset.
| Config | Ranks | Grid | A100 OMP | GFLOP/s | Result |
|---|---|---|---|---|---|
| 8 X100 + 8 A100 × 1 | 8 + 8 | 4×4 | 1 | 57.5 | PASSED |
| 8 + 4 × 2 | 8 + 4 | 3×4 | 2 | 56.2 | PASSED |
| 8 + 2 × 4 | 8 + 2 | 2×5 | 4 | 55.5 | PASSED |
| 8 + 1 × 8 | 8 + 1 | 3×3 | 8 | 47.6 | PASSED |
| X100 only (ref.) | 8 | 2×4 | — | 53.2 | PASSED |
(Earlier ZVL256B-everywhere hetero peaked at 51.3 GF with 8+2×4; A100 ZVL1024B flips the optimum to equal 16 ranks.)
IME2 / llama.cpp and GPU numbers still to come — same methodology as RV2 / F3.
Related
- Banana Pi BPI-SM10 — product page (sponsor)
- Sponsors — how we fund extra RAM SKUs and CI time
- opensolvers/benchmarks