Skip to the content.

The Banana Pi BPI-SM10 is a Pico-ITX board around the SpacemiT K3 SoC (CoM260 / COM3K316128…): 8× X100 general-purpose cores (RVA23, RVV 1.0, VLEN=256) plus 8× A100 AI cores (IME2, VLEN=1024, up to 60 TOPS), integrated IMG PowerVR BXM-4-64 GPU (Vulkan 1.3 / OpenCL 3.0). Ours runs Bianbu 4.0.6 with EESSI on dev.eessi.io/riscv.

Sponsored by Banana Pi — this board was provided by Banana Pi. Product page: BPI-SM10 (K3-CoM260) · docs.

X100 vs A100

K3 is two asymmetric RISC-V clusters. Both implement RVV 1.0 (A100 is not scalar-only); the AI cores trade general-purpose behaviour for a wider vector unit and IME2. Device-tree: spacemit,x100 (cpu@0–7) vs spacemit,a100 (cpu@8–15).

  X100 (cores 0–7) A100 (cores 8–15)
Role General-purpose CPU AI / matrix (IME2)
Clock (typ.) up to 2.4 GHz up to 2.0 GHz
Clusters 0–1 (cluster_cpus 0–3, 4–7) 2–3 (cluster_cpus 8–11, 12–15)
ISA (hart) rv64imafdcvh + bitmanip / vector crypto rv64imafdcv (no h) + same vector ext. family
Hypervisor yes (h) no
RVV yes — VLEN=256 (vlenb=32, measured) yes — VLEN=1024 (vlenb=128, measured on cpu8)
Custom — IME2 · up to 60 TOPS (vendor matrix path)
Default Linux scheduling yes — normal SMP fenced — online in sysfs, but not in a normal task’s affinity mask

Plain RVV on A100 is often slower than on X100 without vendor IME2 code — wide VLEN is there for the matrix path, not as a drop-in “faster X60.”

Getting onto A100 (/proc/set_ai_thread)

SpacemiT’s kernel keeps general threads on X100 even when A100s are idle. Stock taskset -c 8 fails with EINVAL until the thread is registered as an AI thread:

echo $$ > /proc/set_ai_thread   # world-writable on Bianbu; unlocks cores 8–15 for this TID
taskset -c 8-15 …               # now affinity works; process stays on A100 for its lifetime

On our board /proc/set_ai_thread is mode 0222 (write-only). After the write, Cpus_allowed_list becomes 8–15 (X100 is no longer available to that task).

VLEN pitfall: do not start on X100 (VLEN=256), let ld.so cache vector decisions, then migrate to A100 (VLEN=1024) — that pattern can SIGSEGV. Prefer set_ai_thread before exec of the real binary (same pattern as k3_taskset / SpacemiT’s AI runtimes).

For mixed MPI jobs (e.g. HPL), ranks on A100 each echo $TID > /proc/set_ai_thread then pin to cores 8–15 before exec. Equal rank counts are a poor fit for HPL — see heterogeneous balancing below.

Compute paths on this board

All four paths we care about are present — and the Custom / GPU story differs from K1 (X60 IME1 / BXE-2-32).

Compute backends on the Banana Pi BPI-SM10 (SpacemiT K3)

Path On SM10 / K3
Scalar X100 — rv64gc baseline under RVA23
Vector X100 RVV 1.0 · Zvl256b (same VLEN class as X60)
Custom A100 cluster — IME2 matrix path · 60 TOPS (VLEN=1024)
GPU BXM-4-64 — the SKU SpaceMiT’s proprietary OpenCL/Vulkan DDK actually targets (unlike K1’s BXE)

Per-board diagrams for the other machines: VisionFive 2 · Orange Pi RV2 · Banana Pi F3. Homepage logo is the combined view.

Status

Board is on the desk and reachable over SSH with CVMFS + EESSI (EESSI_VERSION_OVERRIDE=2025.06-001, software tree riscv64/generic). eessi_archdetect reports SpacemiT X100. Default user threads only see cores 0–7; A100 work needs /proc/set_ai_thread.

HPL (OpenBLAS RVV)

See also the HPL app overview.

EESSI foss/2025b + HPL/2.3, FlexiBLAS → local OpenBLAS 0.3.34. Two RVV targets matter on K3:

OMP_NUM_THREADS=1 unless noted, NB=192, residual checks PASSED.

Single-cluster baselines (N=12000, 8 ranks, 2×4)

Cluster OpenBLAS target GFLOP/s Result
X100 only ZVL256B 52.0 PASSED
A100 only ZVL256B (VLEN-safe) 15.0 PASSED
A100 only ZVL1024B 35.9 PASSED

A100 ZVL1024B is ~2.4× the portable-ZVL256B A100 HPL rate (~1.45× slower than X100, vs ~3.5× before). Stock NETLIB on A100 was ~0.8 GFLOP/s @ N=4000.

A100 ZVL1024B at smaller N: 22.5 GF @ N=4000, 30.4 GF @ N=8000 (both PASSED).

Equal 16-rank mixed (8+8), same BLAS everywhere

Ranks 0–7 on X100, 8–15 on A100 (set_ai_thread), 4×4 grid:

N BLAS on A100 GFLOP/s Result
4000 ZVL256B 16.2 PASSED
8000 ZVL256B 23.7 PASSED
12000 ZVL256B 27.5 PASSED
12000 ZVL1024B (X100 still ZVL256B) 57.5 PASSED

With ZVL256B on both clusters, equal ranks stall on slow A100 work. With ZVL1024B on A100, equal 8+8 becomes the best mixed config and beats X100-only.

Heterogeneous X100 + A100 (N=12000)

Per-rank FlexiBLAS: X100 → OPENBLAS_DUAL (ZVL256B), A100 → OPENBLAS_ZVL1024B. Wrap: HPL_NX100, HPL_A100_THREADS, set_ai_thread + taskset.

Config Ranks Grid A100 OMP GFLOP/s Result
8 X100 + 8 A100 × 1 8 + 8 4×4 1 57.5 PASSED
8 + 4 × 2 8 + 4 3×4 2 56.2 PASSED
8 + 2 × 4 8 + 2 2×5 4 55.5 PASSED
8 + 1 × 8 8 + 1 3×3 8 47.6 PASSED
X100 only (ref.) 8 2×4 — 53.2 PASSED

(Earlier ZVL256B-everywhere hetero peaked at 51.3 GF with 8+2×4; A100 ZVL1024B flips the optimum to equal 16 ranks.)

IME2 / llama.cpp and GPU numbers still to come — same methodology as RV2 / F3.